Some of the things I've built

Lab

I build tools to explore new ideas and improve ways of working.

They're practical, hands-on, and solve real problems. Most started as something I wanted to do faster or better or turning an idea into a working prototype.

Along the way, they've helped me learn faster, reduce repetitive work, improve collaboration, and streamline day-to-day processes. Some remain internal tools that support my own workflow, while others evolve into prototypes, products, or new ways of working.

Roadmapping Center
Product

Roadmapping Center

Roadmap Command Center translates raw research into prioritized product bets through an evidence-first workflow. Opportunities are scored with a transparent rules engine, organized into a Now/Next/Later board, and linked directly to transcripts, quotes, notes, and decision records for full traceability. An optional AI copilot drafts suggestions and weekly reviews, while final scoring and roadmap changes remain human-controlled for reliability and auditability.

Design-system-aware AI-to-code
Automation

Design-system-aware AI-to-code (with MCP)

Most AI design-to-code tools generate throwaway markup that ignores your component library. I built the opposite: a workflow where the design system is the target language. Every design element must resolve to a real library component, or nothing generates. It uses a Model Context Protocol (MCP) server to give the AI full design context (not screenshots), plus a governed design-to-component mapping layer as the source of truth.

Component status dashboard
Research

UX Heuristic Review Agent

A rules-based UX audit system that automatically evaluates designs against a defined set of usability heuristics, then returns structured feedback with severity explanation, and clear fixes. It makes UX quality assurance more accessible and scalable by converting subjective feedback into a repeatable scoring model, so temas can track improvements over time. In short, it helps us prioritize what to fix first by combining severity, confidence levels and a usability score with actionable insights.

Featured build

Repository AI

An AI-powered research intelligence platform that turns fragmented customer knowledge into connected product intelligence. It links research, designs, decisions, and product strategy into a living knowledge map, helping teams find evidence faster and make more confident decisions.

Every design element must resolve to a real library component, or nothing generates. It uses a Model Context Protocol (MCP) server to give the AI full design context (not screenshots), plus a governed design-to-component mapping layer as the source of truth.

Closed Alpha: Currently being tested with a select group of users.

Check it out →
Design-to-code pipeline
Data

Feature Prioritisation Tool

A self-serve tool (built in R and Shiny) where a team uploads their data as a CSV and it runs the analysis and returns prioritisation recommendations automatically, packaged so anyone can use it, not just analysts.

It started as a Kano model, but people wanted to fold in their RICE scores too, and the interesting problem was that Kano and RICE measure different things (Kano: what kind of value for users; RICE:effort vs. impact). Combining them meant designing a weighting system to reconcile two different lenses, not just bolting one onto the other.

Data

Analytics / Behavioural Automation

A bespoke analytics automation tool (built with GA, R and Shiny) that combined behavioural data, user research and statistical analysis to identify high-impact product opportunities beyond standard GA reporting.

The goal was to uncover where product changes would have the greatest experience and commercial impact, to support opportunity sizing. It helped us priorituse product improvements and highest potential areas based on customer impact.

AI

LLM Evaluation Framework

A lightweight evaluation framework to assess groundedness, factual accuracy, consistency, and task success. The goal was helping determine when an LLM is reliable enough for real users.

The challenge was improving confidence in the product using the model, rather than improving the model itself. It was about measuring the model's output in context by introducing repeatable evaluation criteria, teams can compare prompts, retrieval strategies, and model versions to see what works best for their use case.

AI

Multi-Agent Workflow Experiments

Designing AI teams to solve problems. I explored how to structure multi-agent workflows to solve complex problems, and how to design the right prompts and context for each agent.

Research

Real-World Smart Home Research Lab

How do we evaluate smart-home experiences when the home itself is part of the product? I wanted to usability test smart-home devices in their realistic context, so I built a research lab in a real home. It was designed to be a comfortable, lived-in environment where participants could interact with devices naturally, while still allowing us to observe and collect data.

Research

Listening to Front-line Workers

A scraper that pulls conversations from logistics and warehouse communities on Reddit (via the Reddit API), a way to hear what frontline and warehouse workers actually talk about, in their own words, at a scale interviews can't reach.

Passive research to surface the real problems, language and frustrations of a group that rarely makes it into a research session.