Tinton's ability to translate complex business requirements into intuitive, accessible experiences is exceptional. He doesn't just design screens, he architects outcomes.
Design that
moves the
needle.
Senior UX Manager and Principal UX Designer with 17+ years of leadership experience across India and the UK. I translate complex business challenges into measurable, human-centred outcomes.
Leadership Philosophy
Strategy before pixels. I lead design teams to ask the right question before designing the right answer.
My practice centres on aligning design systems with business roadmaps, empowering cross-functional teams, and building a measurable design culture, where every UX decision is tied to a business outcome.
Strategic Alignment
Every design decision mapped to OKRs and business KPIs, so every pixel earns its place.
Cross-Functional
Bridging product, engineering, and research to create shared ownership of UX.
ROI-Driven
Every project closes with hard data: adoption rates, time savings, or revenue impact.
Team Growth
Mentoring juniors into senior roles through structured rubrics and career frameworks.
What colleagues say.
Reflections from product leaders, engineering managers, and clients I've partnered with across three continents.
Leading a cross-vendor team is never easy, but Tinton brought clarity and vision. The Design for Delight approach genuinely transformed our digital health experience at Boots.
Tinton is that rare UX leader who speaks both design and business fluently. His analytics dashboard UX at AXA XL gave our underwriters a real competitive edge.
Ready to move the needle together?
Open to Senior UX Manager, Principal Designer, and Head of Design opportunities, plus advisory engagements.
Results-first case studies.
Structured using the Minto Pyramid: lead with the outcome, then unpack the context, strategy, and execution.
Co-architected agentic AI experience frameworks across core Intuit products, establishing scalable design patterns that reduced user task completion time by 58%.
Partnered with Intuit UX leadership to guide my team in scaling an agentic UX framework from Intuit Assist in QuickBooks to the enterprise products — establishing trust, control, and transparency patterns that allowed finance teams to delegate multi-step workflows with full oversight.
Created the Vibe Designing framework: AI-native UX that ships working prototypes, not just Figma files.
In partnership with Intuit UX leadership, I coached my team to adopt Vibe Designing — an AI-native framework using NotebookLM, Claude, and Figma AI to transform research and PRDs into traceable, dev-ready code rather than static mockups.
Led a six-person, two-vendor team to an award-winning platform that cut manual inspection dependency by 32%.
UX lead and project manager for Network Rail's Virtual Lineside Inspection platform: owned strategy, research, design direction and the release plan; managed three designers, two BAs and an SME across ten months; won the Digital Transformation Award 2023.
Lifted online quote conversion by 41% for the largest auto insurer in Massachusetts.
Led a three-person team through a six-month, deadline-bound redesign of MAPFRE's legacy Online Bind. Replaced a lengthy single-page form with a guided wizard, extended the brand system with WCAG AA components, and took usability-test completion to 100% error-free.
Led a nine-person team to unify NHS care and commerce in one app, cutting drop-off by 34%.
UX lead and design manager for Boots' personal-healthcare app strategy with the NHS: ran a double-diamond programme across research, service design and delivery, managed four designers and two researchers, and shipped WCAG 2.1 AA-compliant prescription, pharmacy-service and Health Hub journeys.
Led a six-person team to a shipped parking app in a three-month sprint.
Lead Experience Strategist and design lead for KOKO's zero-to-one mobile app: owned strategy, research and design direction, ran stakeholder workshops, coordinated four developers across three releases, and shipped to Google Play with 10K+ downloads.
Led a three-person team from idea pitch to MVP in eight weeks.
UX lead for a social decision-making app: owned the research plan, product strategy and design direction, ran the stakeholder pitch, and coached a BA and a visual designer through a lean, evidence-led sprint to a testable MVP.
Leadership artifacts.
Beyond wireframes: the systems, frameworks, and cultural tools that scale design organisations.
Design System Architecture
Unified component library across IES, PPM, and Compliance. 30% component reuse, 15% faster dev handoff.
Hiring Rubric & Career Framework
Skills-based hiring rubric evaluating UX candidates across 5 competency levels, from IC to Principal.
Team Maturity Model
Quarter-on-quarter capability assessment. 3 junior designers progressed to senior roles under this framework.
Research Operations
Institutionalised Intuit's Follow Me Home practice across 2 vendor orgs, integrated into sprint cycles.
Accessibility Playbook
WCAG 2.1 AA guide adopted across Boots' 3 digital product teams. 100% compliance for medical e-commerce.
Stakeholder Workshops
Design-thinking workshop templates that cut project kickoff ambiguity and turnaround time by 30%.
17 years of enterprise UX leadership.
Spanning FinTech, HealthTech, InsureTech, Government, and Defence sectors across India, the UK, and global clients.
- Lead an 8-member cross-vendor team across IES, PPM, and Compliance platforms.
- Reduced design rework cycle time by 15% by integrating quant and qual insights into the design pipeline.
- Implemented a unified design framework: 25% efficiency boost and 30% faster turnaround within 6 months.
- Applied Intuit's Design for Delight and Follow Me Home methodologies across multi-product environments.
- Led a team of 6 designers for the UK's largest rail network, building inspection apps for engineers.
- Streamlined design-to-development handoff by 15% within 5 weeks, with 30% component reuse.
- Implemented analytics solutions decreasing manual inspection dependency by 32% within 6 months.
- Achieved 66% adoption among inspection engineers within 3 months of rollout.
- Spearheaded UX for a unified e-commerce + NHS healthcare experience for 13M+ customers.
- Increased self-service adoption by 27% through improved UX for consultations, prescriptions, and Health Hub.
- Reduced drop-off rates by 34% through improved navigation and contextual guidance.
- Achieved 100% WCAG 2.1 AA compliance across medical e-commerce and health consultation flows.
- Oversaw the full UX lifecycle for a next-gen analytics platform for underwriters at the UK's largest insurer.
- Achieved 37% adoption and decreased manual data entry by 53% within 9 months of rollout.
- Reduced underwriting analysis time by 15% per case through optimised data hierarchy and visualisations.
- Delivered UX for financial services and enterprise clients in EY's digital transformation practice.
- Collaborated with global stakeholders on complex data-driven products in banking and insurance.
- Pramati: enterprise SaaS product design and B2B interaction design.
- Zen: UX for defence simulation and training platforms, applying human factors in high-stakes contexts.
- Engineering role bridging technical implementation and early interest in user-interface design.
Academic foundations.
Skills & expertise.
17 years of honed capability across UX leadership, research, design systems, and cutting-edge tooling, including Agentic AI and XR.
Skills by category
UX Leadership
UX Practice
Tools & Tech
Industries & domains.
Tools I use daily.
Certifications & credentials.
A commitment to continuous learning, from classical usability methods to cutting-edge XR and AI design.
Certified & credentialed.
Professional Certificate in Data Science & AI (DSAI) for Managers
UX & Interaction Design for AR/VR/MR/XR
Certified Usability Analyst (CUA)
Gen AI Fundamentals
Certified Project Manager
M.Des Industrial Design · GPA 8.5
B.Tech Applied Electronics
Two decades of learning.
Writing & insights.
Perspectives on UX leadership, design strategy, accessibility, and AI-augmented design practice.
Latest articles
Why UX Managers Must Think in OKRs, Not Wireframes
The shift from IC to UX manager is not a promotion, it's a role change. Here's how I learned to lead with outcomes instead of outputs.
Building WCAG 2.1 AA Into Your Design System From Day One
Accessibility is not a checklist, it's a design culture. How we achieved 100% WCAG compliance on Boots' healthcare platform.
Agentic AI in Enterprise UX: Opportunity or Threat?
With Gen AI reshaping how products are built, what does it mean for the future of UX leadership? My take after earning the XR Design cert.
The Team Maturity Model That Grew 3 Junior Designers Into Seniors
A practical walkthrough of the maturity model framework I built at Intuit and how it created a culture of structured growth.
Minto Pyramid for UX Case Studies: Lead With the Answer
Why structuring your case studies with the result first wins stakeholder attention in the first 60 seconds.
How We Achieved 30% Component Reuse at Network Rail in 5 Weeks
The tactical decisions behind building a cross-platform design system for a safety-critical product under tight constraints.
Let's build something
that moves the needle.
Open to Senior UX Manager, Principal Designer, and Head of Design roles. Also available for advisory engagements and design leadership consulting.
Contact details and message form
Leading a two-vendor team to an award-winning inspection platform.
Virtual Lineside Inspection is a desktop platform that lets Network Rail engineers inspect track corridors for vegetation encroachment from the office, replacing hazardous physical surveys across a 20,000-mile network. I led the ten-month engagement as UX lead and project manager: set the strategy and research plan, directed a team of three designers, two business analysts and a subject-matter expert, ran stakeholder workshops and release governance, and delivered a platform that won the Digital Transformation Award 2023.
Summary & context.
The brief was to envision a desktop application that would let railway inspection engineers carry out virtual checks on unmanned, remote stretches of track to monitor vegetation encroachment, all from the safety of an office environment.
The platform combines two data streams: LiDAR (Light Detection and Ranging) survey data of the physical site and forward-facing video footage recorded from the driver's cab of a passing train. Engineers can compare these two sources frame by frame, assess the level of encroachment at any milepost, and generate maintenance work orders, without leaving their desks.
A machine-learning layer highlights vegetation inside the inspection zone on the video feed and flags encroachment zones in the LiDAR view; based on the engineer's confirmation it recommends next actions and routes the work order to the right maintenance crew.
Why manual inspection couldn't scale.
Vegetation creeping into the rail corridor is a persistent safety threat. It blocks the driver's line of sight, contributes to leaf-fall hazards in autumn, and can cause direct track obstructions if left unchecked. With more than 20,000 miles of track spanning the UK, physically walking every section was never going to be sustainable.
The existing process relied on on-foot engineers carrying handheld tablets, conducting periodic surveys, raising work orders, and then handing those over to maintenance crews. It was time-intensive, weather-dependent, and limited by network connectivity in remote areas.
The client needed a way to digitalise the inspection workflow — moving the eyes-on-track activity into an office environment, while keeping the rigour of a trained inspector's judgement at the centre of the process.
What I owned as the lead.
A safety-critical client, a regulated environment, six people across two vendors, and a ten-month programme. Leadership here meant building shared understanding early, protecting the team's focus from a long requirements list, and making every design decision defensible to engineers whose job is preventing derailments.
Strategy & programme plan
Set the ten-month plan, the MVP boundary and three prioritised releases; owned the UX budget and vendor resourcing.
Research leadership
Designed the discovery approach, led field visits with on-foot inspectors, and ran synthesis so the whole team owned the insights.
Stakeholder & vendor management
Facilitated workshops with SMEs, BAs and client leads; aligned two vendors on one process; ran demos and sign-offs each release.
Team development
Coached three designers through critique and pairing; set the design system and definition of done that let them work in parallel.
How the team ran
- Twice-weekly cross-team design review: each designer owned a user group, everyone reviewed everyone.
- Weekly client steering with a one-page status: decisions, risks, asks.
- Fortnightly SME clinic so domain questions never blocked a sprint.
- Release retros feeding user-story reprioritisation for the next drop.
- Needs / Wants / Desires as the single scoping vocabulary across vendors and client.
- One design library, contrast-validated to AA–AAA before any screen was built.
- WCAG-annotated prototypes as the hand-off standard, so accessibility shipped by default.
- Every layout decision traced to a field observation or a workshop output.
Workshops & field visits.
We opened the project with a structured series of workshops involving SMEs, business analysts, and key stakeholders. The goals were practical: surface tacit knowledge, expose current limitations, agree on priorities, and build consensus around an MVP scope.
To ground the team in real conditions, we organised field visits with on-foot inspectors. Walking alongside engineers as they ran live surveys gave us first-hand insight into the constraints of the existing process and what would need to translate into the digital experience.
Field observations
- Inspectors carry tablets to record on-site assessments in real time.
- Coverage per shift is limited by weather and physical access.
- Low or no mobile connectivity in remote sections delays sync with central systems.
- A single faulty hand-held device can derail an entire inspection target.
Needs, wants & desires.
Synthesising workshop output, we organised the platform's requirements into three tiers, sorted by criticality. This gave the team a clear ranking for MVP scope and a shared vocabulary for trade-off conversations with stakeholders.
Needs
- Visual assessment of third-party vegetation where it poses a risk to the railway.
- Annual review of inspection plans and frequencies to ensure tree-risk is adequately controlled.
- Assess risk across immediate-action, action, and alert zones.
- Escalation path to Control when unsafe situations are detected during inspection.
Wants
- Record inspection output against every eighth of a mile, both sides of track.
- Attach digital photos to support each finding and locate follow-on work.
- Assign predefined responses based on inspection findings.
- Record hazardous trees with derailment or harm potential.
Desires
- Flag legally protected or nationally significant trees in the corridor.
- Tag unique tree IDs where previous identifiers don't exist.
- Leaf-fall severity inspections across operational lines in autumn.
- Smart route planning to avoid repeat visits to the same location.
Personas & journey design.
Designing for a data-heavy enterprise platform is closer to sculpting than illustrating. You start with an enormous block of available information and carve away until only what's useful for the task remains. Strip too much, and the tool becomes useless. Keep too much, and it becomes unworkable. The skill is in the balance.
From the workshop and field data we synthesised two distinct user personas, each with their own pain points, goals, and contextual constraints. These gave us a north star for every interaction decision that followed.
Task flow & user journey
With personas in place, we mapped a pain-points and opportunities chart to anchor the ideation, then visualised the ideal task flow for each user group.
Decisions & trade-offs I made
- Side-by-side LiDAR and video as the core canvas. Engineers judge encroachment by comparing sources; splitting them across screens tested as slower and less trusted. It became the most-loved feature.
- Desktop-first, not mobile. The whole point was moving inspection off the trackside into the office; a large canvas for dense data beat the appeal of a companion app.
- Machine learning recommends, the engineer decides. Kept accountability with a qualified inspector, which was the condition for safety sign-off and for user trust.
- Four layouts explored, one chosen with stakeholders. Making the alternatives visible turned a taste debate into an evidence-based decision the client owned.
- Deferred protected-tree tagging and route planning. Real value, but not on the path to proving virtual inspection. Roadmapped as Desires with a date, which kept SMEs on side.
Risks I planned for
Safety-critical decisions on screen
Explicit escalation path to Control, clear confidence indicators on ML recommendations, and audit trail on every work order.
Two vendors, one product
Shared design library, shared definition of done and cross-review of each other's flows so the seams never showed to the user.
Adoption by experienced field engineers
Involved inspectors from the first field visit through every usability round; their language, not ours, in the UI.
Data density overwhelming users
Progressive disclosure per milepost and one primary action per view; caught and tightened after round-one testing.
Branding, layouts & UI anatomy.
Accessibility from day one
All design-library colour combinations were validated to AA–AAA contrast standards using third-party tooling, before any UI was committed to.
Text, buttons, and iconography were sized for clear legibility across every screen size the platform could be used on.
Prototypes were annotated against WCAG guidelines, exposing focus order and screen-reader traversal so engineering could implement faithfully.
Branding & style guide
Layout exploration
Before committing to a final structure, we explored four candidate layouts. Each was assessed against the core constraint: showing dense, side-by-side LiDAR and video data without overloading the inspector's working canvas.
Anatomy of the UI
Once layout direction was set, I focused on key screens in low fidelity to test usability and scalability. The structure that emerged could absorb multiple user types and use cases without breaking down.
Hand-off & release cycles.
Once wireframes were signed off, design moved into the development team's hands. The build was split into three prioritised releases, each one tested against user stories, with the feedback feeding directly back into the next iteration.
The platform serves a specific audience that values functionality over visual flourish — and it had to be built within the constraints of an industry-regulated environment. Managing a focused agile team let us move quickly without compromising design rigour or engineering quality.
Usability testing & insights.
Design is a continuous practice — never truly "complete". Once the platform was in users' hands, we set up structured usability studies, synthesising findings through affinity maps, pattern-identification sheets, and insight templates to drive the next iterations.
How I measured success
Before the first release I agreed a scorecard with the client so each drop was judged against the operational problem — unsustainable manual inspection — rather than against features shipped.
| Level | Metric | Why it matters | Result |
|---|---|---|---|
| North star | Share of inspections completed virtually | Proves the operating model, not just the UI | 32% less manual dependency in 6 months |
| Adoption | Active inspection engineers | Experienced users choosing the tool | 66% within 3 months |
| Quality | Defect-categorisation error rate | Better decisions, not just faster ones | 18% → 3% (2023–24) |
| Delivery | Design-to-dev handoff time & component reuse | Two vendors building one product efficiently | 15% faster handoff, 30% reuse |
Key insights
- Some users felt overwhelmed by the density of content per page — too much type led to attention drop-off.
- The checklist-heavy inspection flow occasionally slowed users down, especially during the preference setup.
- Side-by-side comparison of LiDAR and forward-facing video proved to be the platform's most-loved feature.
- Defect categorisation became significantly more consistent once we introduced guided digital inputs.
The business impact.
Within months of rollout, the platform was driving measurable change across both the inspection workflow and the wider maintenance operation.
Defect-categorisation error rates also fell from 18% to 3% between 2023 and 2024, driven by clearer workflow design, sharper classification patterns, and guided digital inputs at the point of capture.
The project went on to win the Digital Transformation Award 2023.
What I'd take into the next project.
VLI was one of the most complex projects I've led recently, and one of the most rewarding. Having time and access to do proper primary research before designing — workshops, field visits, contextual interviews — was the single biggest factor in getting the solution right.
Collaboration is the multiplier
The team owned different user groups, but we reviewed each other's progress twice a week. That cadence kept the product cohesive end-to-end.
See it through the user's eyes
Choices that feel obvious to a designer rarely feel obvious to a user. Testing kept us honest and informed every design decision.
Plan the product, not just the screens
Structured sitemaps and system-requirement charts helped us scope a ruthless MVP and stay focused on what mattered most.
Impact beyond the product
What I'd do differently as a manager
- Bring the engineering vendor into field visits, not only design; shared empathy shortened every later trade-off conversation where it happened.
- Baseline the defect-categorisation rate before release one so the quality story had a clean starting point.
- Stagger designer onboarding across the two vendors by a sprint; a common ramp-up cost the first release some velocity.
- Agree the post-launch product ownership and measurement cadence during the pitch, so "launch and learn" had a named owner from day one.
Tinton led UX for this project from the very start. His expertise extends well beyond design into communication and attention to detail — both of which proved essential throughout the engagement. He was deeply committed to getting every detail right, has a strong sense of layout and white space, and consistently went above and beyond when adopting new workflow processes. He coordinated and ran workshops with energy, took stakeholder feedback on board, and was quick to act on it. His prototypes earned high marks from business stakeholders.
Leading a nine-person team to put NHS care in the Boots app.
Boots serves millions of people a week as the UK's number-one prescription dispenser and its favourite beauty retailer — but its app treated health and shopping as separate worlds. I led the UX strategy and a cross-vendor team to design a personal-healthcare layer in collaboration with NHS services: repeat prescriptions, pharmacy consultations, Health Hub guidance and family care, unified with the commerce experience and built to WCAG 2.1 AA from day one.
A national pharmacy with a retail app.
Boots is a rare hybrid: the UK's largest pharmacy network and its favourite beauty destination in one brand. Its app had grown up on the retail side — Advantage Card, offers, shopping — while healthcare lived in separate flows with different navigation, tone and sign-in. Customers who came to buy skincare and stayed to reorder a prescription had to change worlds mid-journey, and many simply gave up.
The opportunity was strategic as much as experiential. The NHS was actively shifting routine care from GP surgeries into community pharmacy, and Boots already delivered a large portfolio of services on the NHS's behalf. A personal-healthcare layer inside the app — trusted, accessible and connected to NHS systems — could make Boots the place millions of people manage everyday health, not just buy it.
Active Advantage Card members — the largest logged-in health-and-beauty audience in the UK
Active Boots app users, overwhelmingly on the retail side of the experience
Healthcare services delivered on behalf of the NHS and privately — mostly invisible in the app
Sources: Boots UK "Boots at a glance" and "Boots in numbers" corporate fact sheets.
Drop-offs at the seam between care and commerce.
Analytics told a clear story. Over a third of customers who tried to cross from shopping into a health journey abandoned it — at re-authentication, at duplicate forms, at NHS eligibility copy buried three taps deep, and at inconsistent buttons that made the next step unclear. The cost showed up as lost prescriptions nominated to other pharmacies, abandoned service bookings, and support calls for "can't find" and "can't log in".
The regulator raised the stakes. Anything touching medicines and consultations had to meet WCAG 2.1 AA and clinical-governance standards; anything less was a compliance risk, not just a UX flaw.
Two products, one confused customer
- Separate sign-in and identity for shopping and health.
- Different navigation, tone and components per domain.
- NHS eligibility discovered only at the end of a flow.
- Prescription status invisible until a text message arrived.
- No way to manage a parent's or child's care.
One identity, one journey, NHS-grade trust
- Single sign-on and one account across care and commerce.
- Unified navigation with a dedicated Health Hub.
- Eligibility surfaced up front on products and services.
- Live prescription tracking, connected to the pharmacy system.
- Family profiles for carers, with consent built in.
What I owned as the lead.
A nine-person cross-vendor team, a regulated domain, and stakeholders ranging from pharmacists to NHS liaison to legal. My job was to give that group one plan, one vocabulary and one bar for quality — and to protect the customer's trust in every trade-off.
Strategy & programme
Framed the healthcare-layer strategy, ran it as a Double Diamond with three release gates, and owned the UX budget across two vendors.
Research programme
Set the mixed-methods plan, recruited across beauty, prescription and carer segments, and ran synthesis as shared workshops.
Stakeholder & regulatory alignment
Chaired a fortnightly clinical-governance review with pharmacy, NHS liaison and legal so copy and logic were signed off before design, not after.
Team development
Coached four designers to own a journey each, paired the content designer on plain-English clinical copy, and set the definition of done.
How the team ran
- Weekly steering with Head of Digital and Pharmacy Operations: decisions, risks, asks on one page.
- Fortnightly clinical-governance review for every health-facing screen and sentence.
- Twice-weekly cross-journey critique; each designer owned one journey, everyone reviewed all.
- Accessibility audit as a release gate, not a post-launch ticket.
- One design system spanning care and commerce, contrast-validated to AA before use.
- Content-first: clinical copy drafted and approved ahead of layout.
- Every change traced to a research finding, an analytics signal or a test result.
- Shared Miro source of truth: problem, hypotheses, evidence, decisions, open questions.
Run as a Double Diamond.
I structured the programme on the Design Council's Double Diamond: diverge to understand the real problem, converge on the one worth solving, diverge again on solutions, converge on what to ship. The shape kept a large team honest — no solutioning before the problem was defined, and no shipping before the solution was tested.
Listening to two audiences at once.
Beauty shoppers browse and want delight. Prescription users arrive with intent, urgency and often anxiety. Carers manage someone else's health on top of their own. Discovery had to honour all three, so I ran a six-week mixed-methods programme with the two researchers, triangulating behaviour with attitude.
In-depth interviews across beauty, prescription and carer segments, including two screen-reader users
Analytics events mapped to the existing cross-domain journey to locate every drop-off
Stakeholder workshops with pharmacy operations, NHS liaison, marketing and legal
Desk research: the market was moving toward us
The NHS's own direction of travel underpinned the strategy. Community pharmacy was being positioned as the front door for common conditions, and digital prescription management was becoming an expectation rather than a novelty. That gave the business case teeth and the team a north star.
- The NHS later launched Pharmacy First (Jan 2024), delivering around 5.4 million consultations in its first 13 months — validating the pharmacy-as-front-door bet.
- Roughly four in five Pharmacy First consultations result in a medicine being supplied — a journey the app needed to make effortless.
- NHS App prescription tracking subsequently rolled out across 1,300+ Boots pharmacies — the "where is my prescription?" problem we prioritised was a national one.
- Boots dispenses millions of prescription items a week and is the UK's number-one dispenser — the app was the under-used front door to that volume.
Sources: NHS Business Services Authority; The Pharmaceutical Journal; The Pharmacist; Boots UK corporate fact sheets.
What customers told us
- "I don't trust health advice if it shows up next to a 2-for-1 mascara deal." — Sarah, 34, repeat prescriptions
- "I keep getting logged out when I add a vitamin to my basket from the Health Hub." — Daniel, 51
- "I gave up on the flu jab booking. It asked for the same information three times." — Priya, 28
- "The medicine label text is too small. I take a photo and zoom in." — Margaret, 67, low vision
- "I manage Mum's prescriptions and mine. The app thinks I'm one person." — Asha, 42, carer
- "I trust Boots more than a random pharmacy app. But the app doesn't feel like the shop." — Marcus, 39
Three archetypes and one problem statement.
Synthesis ran as a two-day workshop with the whole team and the pharmacy lead. Three behavioural archetypes emerged, each with a different trust threshold — and a problem statement the team could hold in one sentence.
People who already trust Boots with their health in store abandon the app at the point where shopping meets care, because the app makes them prove who they are, repeat what they've said, and guess what the NHS will cover — right when they most need reassurance.
Job to be done
Reorder the family's repeats and ask a pharmacist a quick question without booking time.
Frustrations
- Logged out mid-task
- Re-entering NHS number
- Health content under marketing noise
Trust threshold
High — health info must feel clinical and clear.
Job to be done
Manage long-term prescriptions independently, read dosage clearly, book a flu jab.
Frustrations
- Small medicine text
- Modals trapping focus
- Colour-only status
Trust threshold
Critical — an accessibility failure means abandonment.
Job to be done
Handle a parent's prescriptions and appointments alongside her own, with consent in place.
Frustrations
- One account, two people
- Phoning the pharmacy for status
- No reminders for others
Trust threshold
High — needs clear consent and privacy boundaries.
How might we…
…make NHS-grade care feel at home in a retail app?
- One identity, clinical visual restraint in health contexts, plain-English copy approved by pharmacists.
…remove every repeated question from a health journey?
- Auto-fill from one profile, eligibility shown before commitment, status pushed rather than fetched.
…make care manageable for carers and for people with low vision?
- Family profiles with consent, WCAG AA by default, readable medicine information.
Diverge on solutions, converge on what to ship.
With the problem fixed, I opened the team up again: ideation sprints per archetype, service blueprints that mapped app steps to pharmacy operations and NHS systems, and rapid prototypes tested weekly. Then we converged — first on a design system that made accessibility the default, then on a scoped MVP.
Service blueprint: the app is only the front stage
Every health journey depended on back-stage actors — the pharmacy system, NHS eligibility checks, the dispensing team, the store. Blueprinting each journey with pharmacy operations exposed where the app promised something the store couldn't yet keep, and where a single API could remove a whole screen.
Discover
● CuriousLands on a product or a health service from search or Health Hub.
Check
● HesitantIs this covered? Am I eligible? Who is it for?
Commit
● UncertainOrder, book or ask — one clear primary action.
Wait
● AnxiousWhere is it? When can I collect?
Collect & return
● ReassuredCollect in store or receive at home; reminders for next time.
Design system: accessibility as the default
Rather than retrofitting compliance screen by screen, I directed the team to build WCAG 2.1 AA into the tokens and components. Every button, form field and status pattern a designer or engineer picked up inherited compliance automatically — the fastest route to 100% across dozens of templates.
4.5:1 contrast minimum
Every colour pairing validated before approval for product flows.
Never colour alone
Prescription status always paired colour with icon and text.
Keyboard & screen-reader semantics
Focus order, live regions for status updates, descriptive labels throughout.
Touch targets ≥ 44px
Critical for older users and motor impairment in prescription flows.
Readable medicine information
Leaflets reflowed for mobile with adjustable text and plain-English summaries.
Plain-English clinical copy
Content designer paired with pharmacists; reading age targeted at 9–11.
Scoping the MVP
I ran MoSCoW with the Head of Digital, Pharmacy Operations and the engineering lead. The rule: if it didn't remove a repeated question, surface eligibility earlier or make status visible, it waited — however commercially attractive.
Release 1
- Single sign-on across care and commerce
- Repeat prescription reorder & live tracking
- Eligibility shown before commitment
- AA-compliant design system
Release 2
- Pharmacy service triage & booking
- "Ask a pharmacist" async chat
- Health Hub guidance
Release 3
- Family profiles with consent
- Medication reminders
- Contextual health nudges on products
Roadmapped
- Video consultations in-app
- Wearables & health-record import
- Private prescriptions marketplace
Signature journeys
Decisions & trade-offs I made
- Unify identity before unifying visuals. Single sign-on did more for retention than any harmonised header; it went in release one while cosmetic alignment waited.
- Clinical restraint in health contexts. Promotions are suppressed on prescription and consultation screens. Commercial teams pushed back; the trust data won.
- Eligibility before commitment. Showing "NHS-funded / self-pay" on the service card added a data dependency, but removed the most-hated dead end.
- Async "Ask a pharmacist" over video. Cheaper, faster to staff, and what interviewees actually wanted for quick questions. Video roadmapped.
- Family profiles in release three, not one. Highest emotional value, but consent and identity rules needed legal time. Sequenced rather than rushed.
Risks I planned for
Clinical-governance sign-off as a bottleneck
Fortnightly review cadence and content-first drafting so approval ran in parallel with design, not after it.
Health data and privacy
Health data never used for marketing personalisation; consent flows for carers designed with legal and data protection.
Two vendors, one design system
Shared tokens, shared definition of done and cross-review so seams never reached the customer.
Promising what the store can't deliver
Service blueprints with pharmacy operations before any status or timing was shown in the UI.
From whiteboard to wireframe
Before anything touched Figma, the team and I spent two afternoons at the whiteboard with a pharmacist and the engineering lead, sketching the three flows the MVP had to get right: onboarding, medicine tracking and appointment booking. Working at marker speed let us argue about sequence and logic — not pixels — and throw away ideas cheaply.
Two decisions were made standing at this board. The scan → review → set → confirm spine for medicine tracking came from the pharmacist walking us through how a repeat is actually dispensed. And the note at the bottom — "how many people, for you or family members?" — is the moment the carer journey stopped being a "could" and became part of the booking flow's first question.
Provocations: lo-fi wireframes we tested
Before any visual design, the team put deliberately rough Figma wireframes in front of customers and pharmacists. Each one was a provocation — a single hypothesis made tangible enough to argue with.
One home for care and commerce — with promotions switched off in health contexts.
Show "NHS · Free" or "Self-pay" on the card, not on the last step.
A four-step tracker so "where is it?" never needs a phone call.
Carers manage others' care with explicit, revocable consent.
Final UI: delivered screens
The shipped screens keep Boots' navy identity for the shell and use NHS blue only where a service is genuinely NHS-funded, so the badge itself carries meaning. Clinical restraint in health contexts, plain-English copy approved by pharmacists, and AA contrast throughout.
Health Hub home
Repeats due sit above the fold with a single reorder action. Guidance is labelled as NHS-reviewed. No offers appear in this view.
Pharmacist assessment and treatment if needed. Ages 5+.
Private service. Check what your destination needs.
Ages 40+. Walk in or book a slot.
Eligibility-first services
"Who is this for?" comes first so eligibility resolves before a customer commits. NHS-funded services carry the blue badge; private ones show a price.
Live prescription tracking
Status streams from the pharmacy system into a four-step tracker with plain-English timing. Colour is always paired with a label.
Family care with consent
Carers switch between people they look after; consent is visible and revocable, mirroring NHS proxy-access rules.
Test for trust, ship in releases.
Standard usability metrics weren't enough — a health journey can be easy and still not trusted. Alongside task success and time-on-task I added a trust-to-act score: after each task, how comfortable would you be doing this for real, today, for yourself or someone you care for? Three moderated rounds with 24 participants, including screen-reader users, ran across the release gates.
How I measured success
| Level | Metric | Why it matters | Result |
|---|---|---|---|
| North star | Cross-domain journeys completed | Proves care and commerce work as one product | Drop-off down 34% |
| Adoption | Self-service health actions (reorder, book, ask) | Shifts routine care into the app | +27% in 3 months |
| Inclusion | WCAG 2.1 AA conformance across health flows | Regulatory bar and the Margaret persona | 100%; zero criticals at audit |
| Trust | Trust-to-act score (1–7) after task | The value proposition is confidence | +2.3 over three rounds |
| Efficiency | Repeat reorder time vs legacy | The most frequent journey | 2.4× faster |
What changed between rounds
- Moved eligibility from the confirmation step to the service card after 7 of 8 round-one users hesitated at payment.
- Replaced colour-only prescription status with icon + text after screen-reader sessions exposed silent states.
- Rewrote consultation copy with pharmacists to a lower reading age; comprehension errors fell to zero in round three.
- Added a "who is this for?" selector at the start of any health journey once carers surfaced as a distinct segment.
The business impact.
Within three months of phased rollout the unified experience was outperforming the legacy journeys on every metric the steering group had agreed. The clearest signal: customers stopped giving up at the seam between shopping and care.
Qualitative signals moved with the numbers: NPS on prescription and consultation flows rose by 18 points, "can't find / can't log in" support contacts fell by nearly a third, and the annual accessibility audit returned zero critical issues for the first time in the platform's history.
Impact beyond the product
What leading in healthcare taught me.
Healthcare isn't a UX category; it's a constraint that raises the bar on everything. Designing to a regulatory standard forced rigour into every decision — and the resulting discipline made the commercial side of the product stronger too.
Build compliance into the system
Retrofitting WCAG is slow and unreliable. Tokens and components that are accessible by default made 100% conformance a property of the system, not a heroic audit.
Governance is a partner, not a gate
Inviting pharmacy and legal into a fortnightly rhythm turned sign-off from a late blocker into early design input.
Trust is the real currency
Clear language, generous targets and clinical restraint did more for conversion than any promotion. In health, every decision earns or burns trust.
What I'd do differently as a manager
- Recruit carers as a segment from week one; they arrived late in discovery and reshaped release three.
- Instrument the trust-to-act score in production, not only in the lab, so post-launch trust could be tracked as a KPI.
- Bring engineering into the service-blueprint sessions with pharmacy operations; several integration surprises would have surfaced a sprint earlier.
- Agree post-launch ownership of the Health Hub content with the pharmacy team during scoping, so governance survived the programme.
Leading a cross-vendor team is never easy, but Tinton brought clarity and vision from the very first workshop. The Design for Delight approach he introduced genuinely transformed our digital health experience. He treats accessibility not as a checkbox but as a craft, and the team learned a huge amount from how he framed every decision around customer trust.
Re-imagining the Online Bind for the largest auto insurer in Massachusetts.
A complete redesign of MAPFRE's legacy auto-insurance quote and bind platform — replacing a friction-heavy single-page form with a guided wizard experience aligned to modern accessibility and conversion standards.
A legacy platform holding back a market leader.
MAPFRE Insurance is one of the United States' most established carriers, offering coverage across auto, home, motorcycle, watercraft, and commercial lines. In Massachusetts alone, they hold the position of largest private passenger auto insurer, largest homeowners insurer, and largest commercial auto insurer.
Despite this market dominance, their digital experience told a different story. The Online Bind platform — the journey customers used to quote and purchase auto insurance — had aged poorly. Internal customer-satisfaction surveys flagged it as a major friction point, with measurable impact on conversion, usability, and customer satisfaction scores.
MAPFRE's leadership made the call to commission a full redesign: a refreshed visual system, a re-architected quote flow, and full alignment to current accessibility standards. I came in as Principal UX Designer and engagement lead, owning strategy, research, design direction and project management end-to-end.
A client who knew the problem, not yet the answer.
The MAPFRE team had decades of insurance experience — from developers through to risk underwriters and product owners. They knew their existing platform was hurting the business, but they didn't yet have a clear picture of what the new experience needed to look like. My first job was to understand the problem deeply enough to define the solution alongside them.
Adding pressure to the brief: an industry insurance convention in the US was coming up fast, and MAPFRE wanted to debut the new product to potential investors there. That hard deadline shaped every prioritisation call I made. We had to be lean, iterative, and ruthless about cutting complexity.
What I owned as the lead.
A market leader with a dated platform, a three-person delivery team, and an investor showcase that could not move. Leadership here meant setting a plan the client could see, getting every discipline aligned in the first week, and holding the line on scope so the right product shipped on the right date.
Strategy & plan
Worked backwards from the convention date into a five-stage plan with deliverables, validation checkpoints and stakeholder touchpoints, so MAPFRE always knew where the investment was going.
Research & insight
Designed the discovery programme: cross-functional workshops, a competitive scan, a heuristic evaluation, five contextual interviews and a 25-card open sort, synthesised into one persona and one problem to solve.
Stakeholder alignment
Put risk, legal, engineering, BA and insurance experts in one room in week one, aligned business, user and technical success lenses, and co-facilitated the MoSCoW that protected the deadline.
Delivery & quality governance
Ran the engagement as its project manager: scope, milestones, hand-off quality, and WCAG 2.1 AA designed into components before any screen was composed.
How the team ran
- Weekly stakeholder review with a one-page status: decisions made, decisions needed, risks to the convention date.
- Workshop series in week one, then a stage-gate review at the end of each of the five stages.
- Working sessions with the BA and the insurance SME to turn underwriting rules into plain-language questions before they reached the wizard.
- Usability findings reviewed with the product owner the same week, with the "fix before launch / fix in v1.1" call made in the room.
- Every candidate feature ran through the MoSCoW filter; "could we also" requests were parked with a dated note, not refused.
- Component-level accessibility sign-off before any screen layout was composed.
- The existing MAPFRE style guide treated as a constraint, not a suggestion, to maximise component reuse by engineering.
- Three layout candidates stress-tested in the open, so the final choice was defensible rather than a matter of taste.
A five-stage UX process.
Working backwards from the convention deadline, I mapped a high-level UX plan covering five stages — each with its own deliverables, validation checkpoints, and stakeholder touchpoints. This gave the client a transparent view of where their investment was going and where they could expect to weigh in.
Discovery
Workshops, competitive scan, heuristic evaluation, contextual interviews
Strategy
Personas, empathy maps, MoSCoW prioritisation, IA, task and user flows
Design
Visual language alignment, layout exploration, hi-fi screen library
Build
Hi-fi prototype, accessibility annotations, engineering handoff
Validate
Usability testing, affinity mapping, iteration loops
Workshops & competitive scan.
I opened with a structured workshop sequence — bringing technical, legal, risk, design, business analyst, and insurance experts into the same room. Getting everyone aligned upfront saved weeks of asynchronous back-and-forth later. By the end of week one, we had shared agreement on business objectives, technical constraints, and the priority slice for MVP.
Next came the competitive and comparative analysis. I evaluated several direct competitors and a few adjacent insurance providers, focusing on product offerings, online purchase flows, mobile and tablet experiences, and "Request a Quote" patterns.
Observations from the scan
- Most carriers offered helper resources — calculators, glossaries, comparison tools, blog content.
- Customers cite lower premiums, easier comparison, and reduced paperwork as the appeal of buying insurance online.
- Pain points centred on lack of product knowledge, fear of online scams, and weak post-purchase support.
- Customers consistently struggled to understand exactly what their coverage included until a claim moment arrived.
- Provider websites generally performed poorly on mobile — a major gap given customer behaviour trends.
Heuristic evaluation
For the existing MAPFRE Online Bind, I ran a Nielsen-style heuristic evaluation focused on the core quote-to-bind functionality. The goal was to surface the most acute usability flaws before any new design decisions were made.
Five conversations that shaped everything else.
I interviewed five participants ranging from late teens to late sixties, all of whom had purchased some form of online insurance within the previous two years. The mix included working professionals and one recent retiree. Each interview blended open-ended exploration with structured probes around their last purchase experience.
From these conversations, four behavioural themes emerged repeatedly:
Needs
- Information presented in plain, easily-understood language — not industry jargon.
- Easy side-by-side comparison of different policy options and what makes each one different.
- Quick answers to questions, ideally without picking up the phone.
Frustrations
- Provider actions that broke trust — surprise charges, hidden exclusions, confusing renewals.
- Encountering insurance terminology with no in-context explanation.
- Difficulty grasping the financial implications of choosing one policy over another.
Goals
- Confidence that the coverage in their quote actually meets their real needs.
- The best possible coverage at the lowest sustainable monthly premium.
- A clear breakdown of premium, deductible, and what is or isn't covered.
- Multi-product discount bundles where they consolidate with one carrier.
Aligning three lenses of success.
With research synthesised, I worked with stakeholders and product owners to define what success actually looked like — but viewed through three independent lenses. Aligning all three was the prerequisite for everything that came after.
What MAPFRE needed to achieve
- Lift online quote-to-bind conversion across auto insurance flows.
- Reduce dependency on agent-assisted quote completion.
- Modernise the brand's digital presence ahead of the industry convention.
- Improve customer satisfaction scores on the Online Bind journey.
What customers came for
- Get a clear, comparable quote in minutes, not hours.
- Understand what their coverage actually includes — in plain English.
- Feel confident enough to bind a policy without calling an agent.
- Use the journey just as well on a phone as on a laptop.
What engineering needed
- A componentised front-end built on the existing MAPFRE design tokens.
- Clean API contracts for the quote engine and risk-rating service.
- Accessibility patterns baked into reusable components.
- Responsive layouts that worked across all major form factors.
Meet Jamie.
Synthesising interview themes, I built one primary persona — Jamie — capturing the most representative behavioural pattern: a working professional who values clarity, fairness, and self-service, but who has been burned by previous insurance experiences and approaches every new provider with cautious scepticism.
A MoSCoW shaped the MVP.
With a fixed deadline and a finite team, ruthless prioritisation was non-negotiable. I co-facilitated a MoSCoW workshop with the project manager, dev lead, and business analyst to bucket every candidate feature into Must, Should, Could, or Won't (this release).
Core to MVP
- Wizard-style quote flow with clear progress
- Plain-English coverage explanations
- Responsive layout across all breakpoints
- WCAG AA-compliant components
Important, deliverable
- Side-by-side coverage comparison
- In-context glossary tooltips
- Save-and-resume quote drafts
- Multi-policy bundle suggestions
Stretch, if capacity
- Live chat with a licensed agent
- Premium calculator widget
- Personalised recommendations
- Document upload at quote stage
Out of scope
- Native mobile app
- Claims management redesign
- Agent-facing back office
- Multi-language support
Decisions & trade-offs I made
- Wizard over a tidier single page. The quick win was restyling the long form; the research said the form itself was the problem. Chunked steps with visible progress and save-and-resume cost more to build and removed the abandonment trigger.
- Extend the brand, don't reinvent it. A visual refresh was in the brief, but the existing style guide meant reusable components and less front-end effort. I spent the design budget on hierarchy, plain language and accessibility instead.
- One persona, not a set. Five interviews pointed to one dominant pattern. A single, well-evidenced Jamie kept every design review anchored on the same person instead of debating edge cases.
- Accessibility at the component level, first. Reading order, focus order and contrast were signed off per component before screens existed, so AA was never a late audit item.
- Live chat, calculators and document upload pushed to "Could". High appeal, high build cost, low impact on the drop-off problem we were hired to fix. Roadmapped with dates so stakeholders stayed on side.
Risks I planned for
A launch date that could not move
MVP defined in week one, MoSCoW gates on every request, and a stage-gate review at the end of each of the five stages so slippage showed early.
Jargon and hidden exclusions eroding trust
Plain-English coverage explanations, "why we ask" tooltips on every field, and a visible premium, deductible and exclusions breakdown before the bind step.
Three people against a legacy platform
Reuse of MAPFRE's design tokens and style guide, a componentised front-end, and the BA and SME used as force multipliers on content and underwriting logic.
Mobile customers on a desktop-era flow
The competitive scan showed carriers failing on mobile. Responsive behaviour was designed per component and tested in hi-fi across desktop, tablet and phone.
Sitemap, task & user flows.
With personas locked and the feature set agreed, I ran an open card-sort exercise to inform the site map. Participants received twenty-five distinct insurance products on index cards and grouped them however felt most natural — then named those groups themselves. The aggregated patterns translated cleanly into the final IA.
Sitemap
Task flow
To anchor every interaction decision in a real customer journey, I built a task flow tracing the steps a customer like Jamie would take to obtain an auto bind quote.
User flow
From task flow, I expanded into a full user flow — explicitly modelling the obstacles, decision branches, and trust-eroding moments Jamie might encounter, and the design responses to each. The extra time invested at this stage paid back many times over once we entered hi-fi design.
Brand discipline meets accessibility.
MAPFRE already had a documented brand identity. Rather than reinvent it, my role was to extend it forward — keeping the brand recognisable while introducing modern accessibility patterns, responsive component behaviour, and a cleaner information hierarchy. Every wireframe respected the existing style guide to maximise component reuse and minimise downstream UI development effort.
On accessibility, I made sure layout, reading order, focus order, and navigation all met WCAG 2.1 AA at the component level, before any screen-level layouts were composed.
Layout exploration
Before committing to a final structure, I sketched three candidate layouts and stress-tested each against the wizard requirements, mobile responsiveness, and the heuristic gaps identified earlier. The exercise made the trade-offs visible to stakeholders and made the final selection an informed, defensible decision.
Hi-fi screen library
Key screens were taken from low fidelity into refined hi-fi, then re-tested with stakeholders. The final structure flexed across desktop, tablet, and mobile breakpoints without compromising the wizard's logic or the brand's voice.
Before & after.
The single biggest experience shift was rebuilding the quote form as a guided wizard. The legacy version asked for everything on one long page — no progress indication, no context for why a question was being asked, and no save-and-resume. Users abandoned the flow at the first sign of friction.
The new wizard chunks the same data into themed steps, surfaces clear progress, and lets the user save and come back later. Every field is justified to the user with a short, plain-language explanation of why it's needed.
Usability testing & validation.
With a working hi-fi prototype in hand, the next job was to stress-test it. I ran moderated usability sessions over video with four participants who matched the persona profile, with each session focused on three structured objectives:
- First-impression read of the new MAPFRE homepage.
- Discoverability and completion of the "Get an Auto Quote" path through to the final quotation form.
- Friction surfacing — anywhere participants slowed down, hesitated, or asked a clarifying question.
I documented findings in an affinity map, clustered into themes, and prioritised the resulting backlog into "fix before launch" versus "fix in v1.1".
Key insights
- Every participant completed every task — a 100% task-completion outcome.
- Zero errors recorded across all four sessions, against the legacy baseline of significant drop-off.
- Participants explicitly described the new homepage as feeling "trustworthy" — a qualitative win that mattered as much as the conversion lift.
- A small subset still slowed down in the preference-collection step; we tightened the copy and added a "why we ask" tooltip before launch.
How I measured success
Before design started I agreed a small scorecard with the product owner, so the redesign would be judged against the problem MAPFRE commissioned us to fix, a quote journey that taxed conversion, rather than against the number of screens delivered.
| Level | Metric | Why it matters | Signal after launch |
|---|---|---|---|
| North star | Online quote-to-bind conversion | Proves the platform enables growth instead of taxing it | +41% within the first quarter of rollout |
| Self-service | Agent calls about unfinished online quotes | Shows customers can bind without picking up the phone | Dropped sharply post-launch |
| Quality | Task completion and errors in moderated tests | Trust is built or lost inside the wizard | 100% completion, 0 errors across 4 sessions |
| Delivery | Convention debut and launch accessibility audit | The date and the AA commitment were both non-negotiable | Debuted on schedule, 0 critical accessibility issues |
The business impact.
The redesigned Online Bind debuted at the industry convention on schedule. Within the first quarter of full rollout, the numbers told a clear story: every key metric moved in the right direction, and the platform was now an enabler of growth rather than a tax on it.
Customer-satisfaction tracking improved measurably across the post-launch period. Agent call volume related to "I couldn't finish the online quote" dropped sharply — freeing agent capacity for the higher-value conversations they were paid to have.
What this project taught me.
MAPFRE was the project that taught me how powerful a hard deadline can be when paired with disciplined prioritisation. The convention date didn't just compress the schedule — it forced clarity at every decision point. Every "could we also..." conversation ran through the MoSCoW filter, and that single piece of discipline kept the team focused on what actually mattered.
Hard deadlines force clarity
Without the convention date, the project would have ballooned. The fixed end-date kept stakeholders aligned on what was MVP and what was post-launch.
Wizards beat long forms — when justified
The wizard worked here because each step was meaningfully different and the customer could see progress. Wizards aren't universally better — context decides.
Workshops compress weeks into days
Getting risk, legal, tech, and design into the same room upfront eliminated weeks of asynchronous email loops. Front-loaded alignment paid back many times over.
Impact beyond the product
What I'd do differently as a manager
- Bring in a dedicated researcher or a second designer. One UX lead across a six-month redesign kept the team lean to the point of risk.
- Instrument the legacy flow before redesigning it, so the 41% lift had a step-level baseline and not just a top-line conversion number.
- Revisit multi-language support sooner. "Won't for this release" was the right call for the convention, not for the customer base MAPFRE serves.
- Agree the post-convention iteration budget before the debut, so every "fix in v1.1" item had an owner and a date, not just a place on a list.
Tinton ran the MAPFRE Online Bind project end-to-end and brought a calm rigour to it that the whole team noticed. He turned a tangle of legacy issues into a clear plan, kept us focused on the convention deadline, and made sure accessibility and brand discipline were never traded away. The redesign hit every metric we set, and the new wizard flow is now the template the team reaches for on every new product line.
Designing trust into agentic AI for enterprise finance.
A UX strategy and interaction framework for embedding autonomous AI agents into the Intuit Enterprise Suite — letting finance teams delegate complex, multi-step workflows to AI while keeping humans in command of every consequential decision.
From tools to teammates.
Generative AI gave software the ability to answer. Agentic AI gives it the ability to act — to plan a sequence of steps, call tools, make intermediate decisions, and complete a goal on the user's behalf. For enterprise finance software, that shift is transformative and terrifying in equal measure: the upside is enormous productivity, the downside is an autonomous system making consequential financial moves the user didn't fully authorise.
Intuit set out to embed agentic capabilities across the Enterprise Suite so that finance teams could hand off repetitive, multi-step work — reconciliation, anomaly triage, compliance checks, reporting — to AI agents. My mandate was to define the UX strategy and interaction framework that would make these agents genuinely trustworthy: powerful enough to save real time, transparent enough that a controller would stake their name on the output.
This wasn't a single screen. It was a design language for a new category of interaction — one the whole organisation could build against.
Autonomy without abdication.
Early prototypes of agentic features tested badly. Not because the AI was inaccurate — the models were strong — but because users didn't trust what they couldn't see. An agent that silently reconciled a month of transactions and reported "Done" left controllers anxious rather than relieved. They re-checked everything by hand, erasing the time savings entirely.
The core tension was clear: the more autonomous the agent, the less the user understood what it did — and in finance, not understanding is not an option. Auditability, reversibility, and accountability aren't nice-to-haves; they're the job.
The design problem became: how do we let an agent do more, while making the user feel more in control, not less?
What I owned as the lead.
A new category of interaction, a team of eight, and product teams across vendors and time zones waiting for a standard. Leadership here meant owning the strategy and the research programme, turning what we learned into principles other teams could build against, and coaching my designers to work with product, engineering and AI/ML as one team rather than hand off to them.
Strategy & framework
Set the direction with Intuit UX leadership: trust, not accuracy, as the design problem. Authored the eight principles and the five-level autonomy spectrum that became the Suite-wide standard and the acceptance criteria for shipping.
Research programme
Designed and sequenced the programme from attitude to behaviour: the 2025 ideation sprint, 18 contextual interviews and diary studies over five weeks, Wizard-of-Oz sessions on the wireframes, and three rounds of moderated testing with 24 finance professionals.
Stakeholder & governance alignment
Co-owned the guardrails with PM, engineering, security and compliance, and set up the fortnightly governance council where scopes and thresholds are agreed before anything is built.
Coaching & scale
Grew four designers from screen owners into pattern owners. Each ran a scenario end to end, from research script to design-system component, with my review as the quality gate rather than my hand on every flow.
How the team ran
- Fortnightly governance council with PM, engineering, security and compliance: every new agent capability, its scope and its thresholds agreed before build.
- Weekly trust review: research, design and the AI/ML lead walked one scenario's plan, trace and ledger together and logged where confidence dropped.
- Design critique run against the eight principles, so feedback was "this isn't Observable", not "I don't like it".
- Playbook office hours for product teams onboarding their first agent: they built, my team reviewed.
- Scenarios before screens: each scenario named the trigger, the agent's actions, the human checkpoint and the record before any wireframe existed.
- One designer owns one scenario end to end, research script to shipped component. I set the bar; they own the work.
- Every pattern ships into the design system with states, copy guidance and accessibility built in, never as a one-off screen.
- Four shared numbers (delegation confidence, task time, undo rate, exceptions) reported by every agentic feature from day one.
Intuit Assist for QuickBooks: an ideation sprint under real constraints.
Before the framework existed, there was a hard deadline. In April 2025 leadership needed credible options to bring Intuit Assist to QuickBooks Online customers by peak season — while the open-ended conversational experience would not be ready for public beta in time. I facilitated a cross-functional ideation sprint with design, product and engineering to answer one question: what ways could Assist come to life in QBO by peak, given the timeline and constraints we face?
The sprint produced 60+ ideas, four voted themes and eleven sketched concepts. More importantly, it surfaced the principles — guided over open-ended, defined expertise over global claims, uniquely Assist over duplicated paradigms — that later became the trust-and-control framework scaled across the enterprise products.
The brainstorm wall
Quantity first, judgement later. The wall below is the team's own sticky notes, clustered afterwards by the job each idea does for the customer. The clusters became the backbone of the concept sketches and, later, of the pattern library.
Eleven concepts, sketched in the QuickBooks shell
The voted themes were turned into concept sketches inside the QuickBooks shell so stakeholders could judge placement and intent, not visual polish. Seven of them are shown here, recreated at mid fidelity from the sprint’s lo-fi frames; the full set also covered Assist Playground, Invoice Reminder, Website Scrape, Generative Templates and Connect to an Expert.
From sketches to decisions
The monthly business brief became the anchor concept. Working the flow on the whiteboard produced the decisions that shaped the public launch and the private beta — and, in hindsight, the first articulation of the Bounded and Observable principles.
Watching people delegate.
To design delegation to AI, we first studied how finance professionals delegate to people. Over a five-week research cycle, the team ran contextual interviews with 18 controllers, accountants, and finance managers, plus diary studies tracking how and when they handed work to junior staff — and critically, what made them trust the result.
Three trust mechanisms surfaced consistently, and they became the backbone of the whole framework.
Managers trust delegated work when they can see the steps, not just the outcome.
Trust scales with clear boundaries — "you can do X, but check with me before Y."
People delegate freely when they know a mistake is reversible and traceable.
What we heard
- "If I can't explain to an auditor how a number got there, I can't use it." — Controller, mid-market
- "I don't want it to ask me about everything. Just the things that actually matter." — Finance Manager
- "Show me what changed before it's final. Not a wall of logs after." — Senior Accountant
- "The scary part isn't errors. It's errors I didn't know happened." — VP Finance
- "I'd let it do the boring 90%. The 10% with judgement, that's mine." — Controller, enterprise
Research programme
I designed the programme to answer one question — what makes a finance professional comfortable handing consequential work to a system? — and sequenced methods from attitude to behaviour so we weren't designing to opinions alone.
| Method | Who / how many | What it told us |
|---|---|---|
| Cross-functional ideation sprint | Design, PM and engineering; 2 sessions, 60+ ideas, 4 themes, 11 concepts | Which Assist experiences could ship by peak without an open-ended chat — and the principles behind them |
| Contextual interviews | 18 controllers, accountants, finance managers across IES, PPM and Compliance customers | How delegation to people works today: showing work, setting limits, owning undo |
| Diary study | 9 participants, two weeks | Which tasks they'd hand off first and what they re-check regardless |
| Wizard-of-Oz sessions | 12 sessions, researcher playing the agent | Where confidence dropped: silent actions, technical logs, irreversible steps |
| Co-design workshops | 3 workshops with customers, PM and engineering | The autonomy levels people would actually accept per task |
| Log & ticket analysis | 18 months of support tickets and usage logs | The repetitive, multi-step work with the highest time cost |
Key findings
- Trust is earned per task, not per product. The same person would delegate matching fully and refuse to delegate posting at all. Autonomy has to be settable at task level.
- Seeing the plan beats seeing the result. Confidence rose most when participants could inspect and edit intent before anything ran.
- Technical logs destroy trust; plain-language traces build it. "Matched 118 transactions by amount and date" worked; stack-trace-style output did not.
- Reversibility is the licence to try. A visible, time-bound undo made people willing to grant more autonomy on the next run.
- Accountability must stay with a named human. Every participant wanted their name — not "the AI" — on the sign-off, for auditors and for themselves.
- Interruptions should be proportional. Asking about everything was as bad as asking about nothing; friction has to scale with risk.
Four moments where delegation has to earn trust.
Research gave us the roles; scenarios gave the team a shared picture of what "an agent doing finance work" actually looks like, hour by hour. Each scenario names the trigger, what the agent does on its own, where the human must step in, and how the outcome is recorded. They became the test scripts for every round of validation.
Month-end reconciliation
- Trigger
- Bank feed closes for the month; 142 unmatched transactions.
- Agent does
- Matches by amount/date, suggests near-miss matches with reasons, drafts adjusting entries, assigns leftovers to owners.
- Human checkpoint
- Approves any adjusting entry over a set threshold; reviews the run summary.
- Outcome & record
- Close shortened by days; every entry traceable to a source and reversible from the ledger.
Anomaly triage
- Trigger
- Spend on a cost centre jumps 38% week over week.
- Agent does
- Clusters the drivers, pulls the invoices, drafts a note to the budget owner, proposes a hold on recurring charges.
- Human checkpoint
- Confirms before any hold is applied; edits the message before it is sent.
- Outcome & record
- Investigation time cut from an afternoon to minutes; audit trail of what was examined and decided.
Regulatory check before filing
- Trigger
- Quarterly filing due in five days.
- Agent does
- Runs the rule set against the ledger, lists exceptions with the rule and evidence, pre-fills the remediation form.
- Human checkpoint
- Signs off each exception resolution; nothing is submitted by the agent.
- Outcome & record
- Zero late filings in the pilot quarter; exception history available for auditors.
Bulk pricing-rule update
- Trigger
- New promotion affects 1,200 SKUs across three regions.
- Agent does
- Drafts the rule changes, simulates revenue impact, flags conflicts with existing rules.
- Human checkpoint
- Reviews the simulated impact; Finance approves before anything goes live.
- Outcome & record
- Fewer manual steps; conflicts caught before launch instead of after.
Four principles for trustworthy agents.
From the research I distilled four design principles that every agentic feature in the Suite would be measured against. These became the shared language for designers, PMs, and engineers — and the acceptance criteria for shipping.
Observable
The agent's plan and progress are always visible. No black-box actions — the user can watch reasoning unfold in real time.
Bounded
Users set the agent's authority up front. Low-risk actions run freely; consequential ones pause for explicit approval.
Reversible
Every agent action is atomic and undoable. A full audit trail maps each change back to the step that caused it.
Accountable
The human is always the named decision-maker. The agent proposes and executes; the user owns and signs off.
…and four that make them work at scale
Proportional
Friction scales with impact. Low-risk steps run silently; high-impact ones pause for an explicit, informed decision.
Explainable
Every suggestion carries its reason, its sources and its confidence — in plain language a controller would use.
Interruptible
Pause, take over or stop at any moment, mid-run, without losing completed work.
Learnable
Autonomy grows with evidence. The system earns expanded permissions through a visible track record, never by default.
Guardrails: what the agent may never do on its own
Principles describe the experience; guardrails are the rules enforced in product. I co-owned these with PM, engineering, security and compliance, and they ship as configuration every product team inherits.
| Guardrail | Rule in product | Pattern that surfaces it |
|---|---|---|
| Permission scopes | Agents act only within scopes the user holds; no privilege escalation | Scope shown in the plan header |
| Impact thresholds | Monetary, volume and irreversibility thresholds force a confirmation gate | Approval gate with before/after diff |
| No external submission | Filings, payments and external communications are never sent by the agent | "Prepared for you to send" state |
| Reversibility window | Every action undoable for a defined period; bulk undo for a whole run | Action ledger |
| Confidence floor | Below a set confidence the agent asks rather than acts | Pause-and-ask card |
| Data boundaries | Only the data the task needs; nothing leaves the tenant | Sources listed on every explanation |
| Named accountability | A human owner is recorded on every run and every approval | Sign-off with name and timestamp |
| Rate & blast-radius limits | Caps on actions per run and per hour until trust is established | Run summary with limits shown |
The autonomy spectrum.
Rather than a binary "manual vs automated," I defined a five-level autonomy spectrum. Each feature in the Suite is placed on this spectrum based on the risk and reversibility of its actions — and users can dial an agent up or down the scale as their trust grows.
Assist
Agent suggests; user does everything.
Draft
Agent prepares work for user review.
Confirm
Agent acts step-by-step with approvals.
Supervise
Agent runs; pauses only on exceptions.
Delegate
Agent completes autonomously; reports after.
Three signature interaction patterns.
The framework produced dozens of components, but three interaction patterns carry the weight of the whole trust model. Each was validated in usability testing against the four principles.
From the whiteboard to the patterns
Before any pattern was drawn, I ran a whiteboard session with the team in April 2025 to work out the anatomy of an agent task, with payroll as the test intent. We traced one job across five columns: intent, input assets, analysis, process and end assets. Under each column we wrote the moments the design had to account for: user inputs and information, operations, user approval and iteration, human intervention versus automation, and work product, implementation and versions. A conversation layer ran along the bottom, tying the canvas together in a single window rather than a chat box.
Three of those answers became the signature patterns below. Intent and input assets, agreed before anything runs, became the Plan Preview. The split between human intervention and automation inside the process became the Live Trace and the autonomy dial. End assets with versions and iteration became the Reversible Ledger.
- Intent and input assets became the Plan Preview. What the agent will act on, and with what, agreed before it runs.
- Human intervention versus automation became the Live Trace and the autonomy dial. Which steps run on their own, which stop for a person, visible as they happen.
- End assets, versions and iteration became the Reversible Ledger. Every output versioned, traceable and undoable.
The Plan Preview
Before acting, the agent shows its full step-by-step plan. Users can edit, remove, or reorder steps — approving intent before execution.
The Live Trace
As the agent works, each action streams into a live, human-readable trace. Users watch progress and can pause or intervene at any point.
The Reversible Ledger
Every completed action lands in an audit ledger where any single step — or the entire run — can be undone, with a full trail for compliance.
Testing the shape of control first.
Before any visual design, the four core patterns were wireframed at desktop scale and put through Wizard-of-Oz sessions. The question at this stage wasn't "does it look right?" but "does the person feel in command?"
Show intent as editable steps, each tagged auto or confirm.
Progress in plain language with pause and take-over always visible.
High-impact steps stop and show a before/after diff with the reason.
Every step undoable individually or as a whole run; exportable for audit.
The patterns, productised.
Hi-fi versions on the Intuit design system for the month-end reconciliation scenario. The visual language is deliberately calm — the agent is a colleague doing careful work, not a feature shouting for attention.
Plan preview
Four steps, two run automatically, one needs approval. The autonomy level and the reversibility promise are visible before anything runs.
Live trace
Completed, current and queued steps in plain language with confidence shown; the agent pauses on anything it isn't sure about.
Vendor invoice #44871 (Oct 28) has no matching accrual. Amount and vendor match last month's pattern; confidence 0.91. Source: AP ledger, bank feed.
Approval gate
Before/after values, the reason and its sources, and the user's name on the decision. Undo stays available for 30 days.
What shipped: the Intuit Intelligence widget, state by state
The patterns did not stay in samples. They shipped as the Intuit Intelligence widget, one conversational surface that opens as a side panel and expands to an immersive window, with the same states in both sizes. The board below is the final UI state matrix: nine states, from loading to the metering banner, across the small and medium panel and the immersive view. A left rail carries threads, a new conversation and history; the footer carries the generative-AI disclosure on every state.
Scroll sideways to see all nine states in both sizes. Six of the immersive states are enlarged below.
Scoping the first release
Eight principles describe the experience; the first release had to prove them on one scenario, month-end reconciliation in IES, with a team of eight and other product teams waiting. I ran the MoSCoW with the PM and the AI/ML lead against a single test: does this make a controller more willing to delegate on real company data?
Proves the trust model
- Plan preview with editable, re-orderable steps
- Live trace in plain language, pause and take over
- Approval gate with before/after diff and reason
- Reversible ledger with single-step and whole-run undo
Makes it usable daily
- Autonomy dial per task, Assist through Delegate
- Named human sign-off on every run
- Interruptions proportional to impact thresholds
- Confidence shown on every step and suggestion
Grows with evidence
- Autonomy that expands from a visible track record
- Boundaries remembered across products
- Agent-to-agent handoffs on a shared ledger
- Trace export straight into the audit system
Guardrails, not features
- Agent submits a filing or posts a payment on its own
- Silent scheduled runs with no trace
- Permissions beyond the user's own scopes
- Delegate as a default autonomy level
Decisions & trade-offs I made
- Trust, not accuracy, as the design problem. The models were strong and the early prototypes still failed. Pointing the research at delegation behaviour changed what we built: visibility and control before more capability.
- A pattern language, not a feature list. Principles and a spectrum cost more up front and let several product teams build without me designing every flow. A feature list would have shipped faster once and scaled never.
- Plan preview before live trace. Both tested well; seeing intent before execution moved confidence more than watching progress, so it shipped first and set the interaction model for everything after.
- Supervise, not Delegate, as the IES launch level. Full autonomy was technically possible. Launching at a level that pauses on exceptions let the track record, and the trust, build before the dial moved.
- Guardrails as configuration, not guidance. Rules that live in the product are inherited by every team; rules that live in a document are optional. Co-owning them with security and compliance made them enforceable.
Risks I planned for
Silent failure in a financial system
Every action streams into a plain-language trace, lands in a reversible ledger and pauses on uncertainty, so nothing consequential happens unseen. "Errors I didn't know happened" was the fear we designed against.
Over-asking that erases the time saving
Interruptions proportional to impact: thresholds on money, volume and irreversibility decide what runs and what waits for a person. Re-tested against the "don't ask me about everything" complaint every round.
Pattern drift across vendors and time zones
Components with states and copy guidance in the design system, a self-serve adoption playbook, and a council that sees every new capability before build.
Autonomy outrunning evidence
Launch levels set below what is technically possible, autonomy that grows only with a visible track record, and undo rate and exceptions tracked per feature so the dial can move both ways.
Testing for trust, not just usability.
Standard usability metrics weren't enough here — a feature could be easy to use and still not trusted. So alongside task-completion and time-on-task, we introduced a delegation confidence score: after each test task, participants rated how comfortable they'd be letting the agent do that work unsupervised on real company data.
We ran three rounds of moderated testing with 24 finance professionals, iterating the plan-preview and live-trace patterns between each round. Confidence rose sharply once users could see and bound the agent's behaviour, confirming the research thesis: transparency, not accuracy alone, is what unlocks delegation.
Key insights
- Showing the plan before execution raised delegation confidence more than any post-hoc summary.
- Users who set explicit boundaries once were far more willing to grant higher autonomy later.
- The single most-valued control was a prominent, always-available "undo last action."
- Human-readable traces beat technical logs decisively — plain language was a trust multiplier.
How I measured success
Before the first test round I agreed four shared numbers with the PM and the AI/ML lead, and they are the numbers every agentic feature in the Suite reports today. The framework was judged on whether finance professionals actually delegated, not on how many components we shipped.
| Level | Metric | Why it matters | Signal |
|---|---|---|---|
| North star | Time to complete delegated workflows | The saving only exists once people stop re-checking by hand | 58% reduction on delegated workflows |
| Trust | Delegation confidence score after each task | The leading indicator for whether autonomy can grow | Rose each round once plan preview and bounding shipped |
| Quality | Undo rate and exceptions per run | Shows whether the agent is trusted, and whether it deserves to be | Tracked per feature; reviewed by the governance council |
| Adoption | Products building on the standard | Proves a pattern language scales beyond one team | IES live and expanding, PPM live, Compliance in pilot |
From one team's patterns to every product's standard.
Designing the patterns was half the job. Getting multiple product teams, across vendors and time zones, to build agentic features the same way was the other half — and the part that made this a leadership project rather than a design one.
Pattern library in the design system
Plan preview, live trace, approval gate, action ledger, autonomy dial and pause-and-ask shipped as components with states, copy guidance and accessibility built in — 40+ components in total.
Adoption playbook
A "how to add an agent to your product" guide: scenario template, autonomy-level worksheet, guardrail checklist, validation script. Teams self-serve; my team reviews.
Governance council
Fortnightly review with PM, engineering, security and compliance for any new agent capability — the forum where thresholds and scopes are agreed before build.
Shared measurement
Every agentic feature reports the same four numbers: delegation confidence, task time, undo rate, exceptions. Comparable across products from day one.
Where it landed
| Product | First agentic capability | Autonomy level at launch | Status |
|---|---|---|---|
| Intuit Enterprise Suite | Reconciliation & anomaly triage | Supervise (pauses on exceptions) | Live, expanding |
| PPM | Bulk pricing-rule drafting with impact simulation | Confirm (step-by-step approval) | Live |
| Compliance | Pre-filing rule checks and remediation drafting | Draft (prepares, never submits) | Pilot |
| Further products | Onboarding via the playbook | Set per task | In adoption |
The business impact.
The framework moved from prototype to a phased rollout across the Enterprise Suite. Beyond the headline efficiency numbers, the biggest win was behavioural: users stopped re-checking the agent's work by hand, which is where the real time savings finally materialised.
The autonomy spectrum and four principles were adopted as the organisation-wide standard for agentic features, giving every product team a shared, defensible foundation to build on. What began as one framework became the way Intuit designs AI that acts.
What designing for agents taught me.
Agentic AI is the biggest interaction-model shift since the touchscreen, and it rewards the fundamentals rather than replacing them. The winning moves weren't clever prompts or novel widgets — they were the oldest ideas in UX: show your work, respect the user's authority, make mistakes recoverable.
Transparency is the interface
With agents, what the system reveals about its reasoning is the primary UX. Visibility does more for adoption than raw capability.
Autonomy is a dial, not a switch
Letting users grow the agent's authority over time turned scepticism into trust. Nobody delegates everything on day one.
Principles scale better than screens
A shared framework let eight people and multiple product teams build coherent agentic UX without me designing every flow.
Impact beyond the product
What I'd do differently as a manager
- Instrument undo rate and exceptions in the earliest prototypes, so each product's launch level rested on data rather than judgement.
- Bring security and compliance into the sprint, not only the council. Guardrails negotiated after design cost rework that would have been free in week one.
- Pair each product team with one of my designers for their first agent instead of relying on the playbook alone. Self-serve scaled reach; pairing would have scaled quality faster.
- Plan for vendors and time zones from day one. A fortnightly council works for approvals but is too slow for design questions; async design review with a 24-hour turnaround came later than it should have.
Tinton framed the agentic AI problem in a way the whole org could rally around. Instead of designing features, he designed the principles and the spectrum that every team now builds against. He kept the human firmly at the centre of an AI-first product — which is exactly the leadership this moment demands.
Leading a six-person team to a shipped app in three months.
KOKO is a zero-to-one mobile app that helps drivers find and book parking near their destination and earn shopping perks for it. I led the engagement end to end: set the strategy and research plan, ran discovery workshops with the founders, owned design direction, and coordinated a product owner and four developers across three releases to a live app on Google Play.
Less circling, more shopping.
Finding a parking spot in a busy area is rarely quick. Drivers can lose anywhere from a few minutes to an hour hunting for a space that fits their car and sits close enough to where they're going — mostly because they're making decisions without reliable information about what's available.
KOKO asked me to design a mobile app that fixes this for drivers and, at the same time, gives local shopkeepers a reason to sponsor it. The app surfaces free or nearby parking around a destination, lets drivers book a slot in advance, and rewards each booking with discount coupons and redeemable points from participating stores. Target users: anyone who drives.
Of parking-app users drop off out of frustration with existing apps (public forum research)
Turns and 3 full circles observed for a single car before it found a spot
Sides to satisfy at once: drivers who want a spot, shop owners who want footfall
A parking app people wouldn't abandon.
Parking is frustrating enough on its own; using an app on the move that buries the relevant information makes it worse. Forum research showed a meaningful share of drivers give up on parking apps altogether because they're hard to use under pressure.
The client's brief was to envision something more useful than a map of spaces: an app that helps drivers find the right spot while giving store owners a channel to offer discounts that bring those drivers into the shop. On top of that core loop they wanted advance slot booking, round-the-clock camera monitoring of parked vehicles, and valet parking.
Roadmap
A three-month plan running discovery, strategy, design, build and launch-and-learn as overlapping tracks, so development could start on validated flows while later screens were still being explored.

What I owned as the lead.
A startup client, a single designer and four engineers, and a three-month window. Leadership here meant setting a plan the whole team could see, making the hard scoping calls early, and keeping the founders' ambition tethered to what users actually needed.
Strategy & roadmap
Translated the founders' vision into a three-month roadmap with overlapping discovery, design and build tracks and three release gates.
Research & insight
Designed and ran observation, interviews and competitor analysis; converted findings into an insights map the whole team worked from.
Stakeholder management
Facilitated workshops with the founders and PO, negotiated MVP scope against a long feature wish-list, and ran client demos every release.
Delivery & dev coordination
Owned hand-off quality, sat in sprint planning with four developers, and closed the loop from each release's feedback into the next.
How the team ran
- Weekly founder check-in with a one-page status: decisions made, decisions needed, risks.
- Design review with the PO before anything reached engineering.
- Joined developer stand-ups during build; answered UX questions same-day.
- Release retro after each of the three drops, feeding the next backlog.
- User stories prioritised jointly with the PO, then locked per release.
- Annotated hi-fi specs with states and edge cases, so developers weren't guessing.
- A reusable component library so four engineers built one consistent product.
- Every enhancement request routed through the CJM: which pain point does it remove?
Understanding the brand, the market and the driver.
Discovery began with the business: what problem KOKO was solving, what its objectives were, and what made the brand distinct. From there I studied the market, analysed the competition and set out to understand drivers' needs and pain points first-hand.
Observational research
To ground the survey questions in real behaviour I spent time watching high-traffic parking areas — deliberately covering different parking types (garages, multi-level structures, open lots, and street parking). I timed how long drivers took to find a spot and to pay, and noted their driving patterns. One car alone made more than twelve turns and three loops before parking, then queued five minutes to pay.

- Popular areas have very limited parking.
- No way to know a lot has space until you arrive.
- Physical tokens and tickets get damaged or lost.
- Most private lots don't take online payment.
- Few apps allow multi-day bookings or extending time.
- Hard to find your car later if you forget the spot.
- Poor map experiences for locating slots.
- No security for parked vehicles.
Contextual user interviews
To go deeper on the day-to-day problems, I interviewed five private car owners who drive daily to college or the office. The transcripts were rich but scattered — a reminder that raw interview data needs structured synthesis before it yields conclusions.


Competitive analysis
I benchmarked the parking apps already on the market and compared the ones with elements closest to KOKO's ambition. ParkMobile and RingGo came nearest to the target — map views of availability, navigation and reservation. PayByPhone informed pay-for-a-spot and extend-stay; SpotHero the reservation model; Strava's mapping and Sharvy's kudos, badges and stats shaped the community and rewards thinking. Google Maps set the bar for wayfinding.

Goals, features and the right scope.
Strategy set the goals, determined the feature set and estimated the scope that would deliver the best business impact. It combined brainstorming with the research outputs — personas, user stories, an empathy map and a customer journey map.
User insights
Workshops, observation, online research and the interviews together gave a clear picture of how people park, what gets in their way and what they expect. Everything was plotted onto a single insights map.

User groups & persona
City parking is stressful because drivers decide without full knowledge of safety, rules and availability. From the insights I built a primary persona, Pauline: a busy, cautious car owner who often drives downtown and wants safe parking close to her destination, but loses time to street parking and confusing signage.

Customer journey mapping
A current-state journey map exposed every friction point and opportunity, and became the source for the feature enhancements proposed to the KOKO team.

Task listing
Tasks are the atomic units of activity — the taps, drags and scrolls. I listed them abstractly, referencing motions rather than UI widgets, so they could drive flows without pre-deciding the interface.

User flow
I whiteboarded the flows and structure, deliberately giving users more than one route to each goal, each tailored to the high-level journeys uncovered in research.

Information architecture
With the feature list mapped to pain points, a detailed IA laid out where each capability would live — giving the whole team a shared view of how the platform fits together and where to focus.

Scoping the MVP
The founders' brief included advance booking, 24/7 camera monitoring and valet parking on top of the core find-and-book loop. I ran a MoSCoW session with the PO and dev lead to agree what three months could prove, and parked the rest on a dated roadmap rather than saying no.
Core loop
- Find nearby / free parking on a map
- Book a slot in advance
- Online payment
- Rewards: coupons & points on booking
Confidence
- Auto-saved spot & countdown timer
- Extend booking duration
- Booking history
Community
- Report busy / unsafe areas
- Decode parking signs
- Nearby incidents
Roadmapped
- Camera monitoring of vehicles
- Valet parking
- Multi-day bookings
- iOS build
Decisions & trade-offs I made
- Map-first, not list-first. Observation showed the real question is "where, relative to me?" — so availability lives on the map, with the list as a secondary view.
- Rewards framed as parking value, not generic points. Interviewees understood "get your parking back" instantly; abstract loyalty points tested as noise.
- Three releases instead of one big launch. Smaller drops meant real feedback after week five, not week twelve, and gave the founders visible progress for investors.
- Defer camera monitoring and valet. High founder enthusiasm, high build cost, low impact on the drop-off problem we were hired to fix. Roadmapped with a date, which kept trust intact.
- Component library before hi-fi screens. Cost a week up front; saved far more with four developers building in parallel.
Risks I planned for
Trusting a slot you can't see
Show availability count and location clearly; confirm the booking with a clear reference the driver can show at the gate.
Scope creep from a passionate founder
Every new idea mapped to the journey map and the release plan in the weekly check-in — visible, dated, and never lost.
Four developers, one designer
Component library plus annotated specs so engineering never blocked on design; same-day answers in stand-up.
Use on the move
Large touch targets, high-contrast map pins and one primary action per screen, tested on mid-range Android devices.
Comforting, engaging, minimum fuss.
With research and objectives in hand, design focused on getting drivers to their goal as quickly as possible with as little friction as possible.
Branding guidelines & accessibility
Every component was built to be compatible and reusable within the KOKO style guide, with a minimum of WCAG AA accessibility across the application — contrast, touch targets and legibility checked for use on the move.

Exploration & testing
To get ideas flowing I sketched a series of paper prototypes. These lo-fi artefacts made it cheap to test and iterate before committing to screens.


Low-fi wireframing
Key screens were worked up in low fidelity to find the most usable and scalable structure. After testing and refinement, a unified layout emerged that flexed across user types and use cases.
Parking space booking

Offers and points redemption

Three releases, one MVP.
After stakeholder sign-off on the wireframes, designs were handed to the four-developer team building the MVP. Development ran in three releases prioritised by user story; each release was tested and the feedback rolled into the next.
Final UI
Once the lo-fi wireframes had been tested with users and the flow finalised, hi-fi wireframes and visual design were produced and delivered for development.

Delivered experiences
Find safe parking
Ideal spots ranked by distance to the destination, availability and safety.
Decode parking signs
Check the rules for a spot without deciphering confusing signage.
Help the community
Shared parking experiences improve future availability predictions.
Auto-saved parking
Your spot is saved automatically with a countdown based on the vehicle's location.
Incidents nearby
Live parking and road conditions so drivers can route around trouble.
Offers nearby
Relevant store offers surfaced by proximity to where you've parked.
Usability testing and what's next.
Design is never finished. Once real users had the product, we reviewed performance, gathered feedback and improved against both user needs and business goals.
I assembled a prototype to surface usability issues. Even with limited functionality, participants could grasp each screen's structure and flag anything that would hurt the experience. Findings were clustered in an affinity map, prioritised, and fed back into navigation and other areas.
How I measured success
Before the first release I agreed a small scorecard with the PO so each drop could be judged against the problem we were hired to solve — drivers abandoning parking apps out of frustration — rather than against feature count.
| Level | Metric | Why it matters | Signal by release 3 |
|---|---|---|---|
| North star | Bookings completed without support contact | The app works on its own, in the car | Steady rise across the three drops |
| Adoption | Installs & repeat bookings | Proves the value loop for drivers and shops | 10K+ installs on Google Play |
| Usability | Task success & SUS per release | Benchmarkable against industry (68) | Above benchmark by release 3 |
| Delivery | Design defects raised in QA | Hand-off quality with four developers | Falling release over release |
What next?
- Keep testing and iterating to sharpen the design.
- Cross-promotions with local businesses as a revenue stream.
- Hands-free, two-way voice interaction via NLP for a touch-free experience while driving.
- Park with AI: pre-book spots from the user's calendar.
- Smart cars + smart parking: vehicles negotiating a spot with sensors, no human input.
Key learnings
Listening to customers and would-be customers is everything. Understanding how they think and feel made it possible to empathise and design for the problems they actually have. Most of those problems aren't new — plenty of parking apps already ship these features. The point was never to reinvent the wheel but to listen closely enough that the experience becomes obvious and users don't have to think.
And it doesn't end at launch: shipping the MVP to real people is what validates the assumptions and answers the questions users still have. Further iteration towards delight is the path from here.
Impact beyond the product
What I'd do differently as a manager
- Recruit a second designer or a dedicated researcher — one UX/VD lead across four developers is lean to the point of risk.
- Instrument analytics in release one, not release two, so the north-star had a baseline from day one.
- Run the observation study in the client's own cities earlier; local parking behaviour shaped several late changes.
- Agree the post-launch ownership model up front, so "what next" had a team and a budget, not just a list.
From the KOKO team.
I hired Abraham to lead our UX effort and couldn't have been happier. I have found him to be personable, professional, hard working and dedicated to his craft. His work was invaluable to creating a better product. His studies on product usability showed us in concrete terms just how much we misunderstood our customers and how poor our UI design choices were. He showed us where we had gaps in our understanding and guided our team to create a far more user friendly and effective product.
I've had the pleasure of collaborating with Tinton on a couple of projects. Tinton is a passionate, dedicated and talented UX designer with strengths in creative and strategic thinking. He impressed me with how well he connected with engineers and has always been keen to learn more about the technical architecture of the systems. Tinton's work is intuitive and captures the client's needs while being firm on designing the right solution without shortcuts. Keep up the good work.
Leading a lean team from pitch to MVP in eight weeks.
Opatter is a social decision-making app: photograph the item you're hesitating over and get an instant "Awesome" or "Meh" from the five friends who know you best. I led the engagement end to end — owning the research plan, product strategy and design direction, running the stakeholder pitch, and coaching a business analyst and a visual designer to a tested MVP on an eight-week clock.

Too many options, not enough confidence.
Choice overload at the shelf
- Shoppers face so many options that picking one becomes stressful, not satisfying.
- Many describe a "choking" moment of confusion over what they actually came to buy.
- The result: overspending on half-satisfying items, or walking out with nothing.
People relax when a friend is with them
- Shopping with a close friend who knows your taste lowers anxiety and speeds decisions.
- That friend is rarely physically there when the decision happens.
- So: bring the friend into the moment through the phone.
A circle of five, not a crowd of hundreds
- Ask up to five of your closest friends — from your social graph or WhatsApp list.
- Three or more "Awesome" votes means go ahead; otherwise, hold off.
- Fast, binary, trusted — the opposite of scrolling anonymous reviews.
Buy confidence, not just products
- People buy value and satisfaction, not apps.
- Opatter saves time and money with the reassurance you only get from your closest circle.
- Works whether you're 16 or 60.
Design process
Discover
Stakeholder brief, market & competitor scan, observation, interviews
Define
Persona, empathy map, journey, task flow, MoSCoW scope
Design
Lo-fi wireframes, style guide, hi-fi screens & interactions
Build
Feasibility with engineering, spec & hand-off, MVP
Test & learn
Two usability rounds, SUS, decision log, iterate
What I owned as the lead.
This was a small, fast engagement — exactly the kind where leadership shows up as clarity, sequencing and coaching rather than headcount. I set the plan, made the calls, and made sure the team learned while shipping.
Strategy & scope
Framed the problem, wrote the value proposition, and scoped an MVP the team could test in eight weeks — not a feature wish-list.
Research plan & rigour
Designed the research plan, interview guide and participant criteria; ran synthesis workshops so insights were shared, not siloed.
Stakeholder alignment
Built and delivered the idea pitch to leadership; negotiated scope with the BA and feasibility with engineering before design began.
Coaching & craft
Ran twice-weekly critiques, paired the visual designer on hierarchy and components, and set the definition of done for every deliverable.
How the team ran
- Monday plan / Friday demo to stakeholders — nothing waited more than a week for feedback.
- Twice-weekly design critique with a written decision log.
- Research synthesis run as a workshop with the BA and VD, not a designer-only readout.
- Engineering feasibility check before hi-fi on any new interaction.
- One source of truth in Miro: problem, hypotheses, evidence, decisions.
- Every design change traced to a research finding or a test result.
- "Definition of done" per artefact: persona, journey, flow, wireframe, spec.
- Clear ownership: BA owned requirements, VD owned the UI kit, I owned experience decisions.
Friendship matters.
Staying close to a few people keeps us grounded in every part of life. Social media optimises for a large network; decisions that matter lean on a small one. The research set out to quantify that gap and understand how the target users actually stay connected.
On average, a person has…
Social media friends
Close, bonded relationships
People they share their most private details with
Source: Harvard Business Review


Other findings
- 88% of people use their smartphone while shopping in store.
- 4.8 billion people use social media — 59.9% of the world and 92.7% of internet users.
- Across every age group, "friends & family" is the top reason people use social media.
- 79% make purchase decisions based on suggestions from friends and family.
Competitive landscape
I benchmarked how existing tools help people decide while shopping. Most either lean on strangers' reviews or on slow, one-to-one chats. Nobody offered a fast, structured verdict from a small trusted circle.
| Approach | Examples | Strength | Gap Opatter fills |
|---|---|---|---|
| Crowd reviews & ratings | Amazon, Flipkart, Google Shopping | Volume of opinions | Strangers' taste ≠ yours; slow to read in-aisle |
| Messaging a friend | WhatsApp, Messenger | Trusted source | One-to-one, unstructured, easy to miss, no verdict |
| Social polls | Instagram Stories, Facebook polls | Fast, visual | Broadcast to everyone; no response deadline |
| Visual discovery | Pinterest, style apps | Inspiration | Helps browse, not decide on the item in hand |
| Opatter | — | Five trusted friends, timed, binary verdict | Decision in minutes, in the store |
How the target users stay close.
Uncover the motivations and pain points in how Gen Z and Millennials connect with the people closest to them, and how they prioritise and maintain those relationships. Interviews ran 30–40 minutes, in person or by phone.
Ages 11–46, building a full-time career, have felt disconnected from friends, value social relationships, tech-savvy, and use at least one social platform to keep in touch.

Findings
- Millennials are career-focused; they seek out friend activities only when they have time and energy.
- Gen Z stay connected through shared activities; face-to-face wins, but messaging around shared interests bridges gaps.
- Busy schedules mean people need a nudge — a post or an old photo — to reach out.
- Millennials value and want to nurture existing relationships over new ones.
- Over-focus on career brings a background sense of loneliness.
- The longer the silence, the harder it is to reconnect without an icebreaker.
Friends' opinions outweigh ratings & reviews
Published research backs the hunch: the volume of reviews from friends shapes following behaviour far more than the crowd's. Crowd valence and variance showed no effect once friend reviews were ignored — and a negative one when they were included. In short, the influence people attribute to reviews mostly comes from reviews by friends.
Source: ScienceDirect, "Friend reviews vs. crowd reviews" (2018)

Two concerns shared by both groups


Meet John.
Synthesising interviews and research into pain points and opportunities produced a behavioural persona, John, which set the direction for the solution's intent and core needs.

Making a purchase decision is tough for John
An empathy map captured the concerns and confusions running through John's head at the point of purchase.

Customer journey mapping
Plotting John's experience end to end showed exactly where a digital intervention could relieve the most friction.

The Opatter logic


Task flow

Prioritising the MVP
With a fixed eight-week window, I ran a MoSCoW session with the BA and engineering to agree what a testable MVP actually needed. The rule: if a feature didn't help a user reach a verdict faster or trust it more, it waited.
Core loop
- Snap item → send to five friends
- Awesome / Meh verdict with majority rule
- Response-time control
- Social sign-in & friend selection
Trust & speed
- Visible countdown to verdict
- Rubber-band gesture (with slider fallback)
- Result history
Delight
- Comments on a verdict
- Gift mode for buying for others
- Streaks / friend badges
Out of scope
- In-app purchasing
- Retailer integrations
- Public / crowd voting
- iOS build
Decisions & trade-offs I made
- Cap the circle at five. Research showed decisions lean on 2–5 close ties; larger groups in testing produced split, slower results. Fewer voters, faster confidence.
- Binary verdict, not a rating. A 1–5 scale invited hedging. "Awesome / Meh" gave a clear majority in one tap and mirrored how friends actually talk.
- Ship the gesture, keep the slider. The rubber-band drag tested as delightful, but not universally usable — so it launched alongside an accessible slider rather than replacing it.
- Post to friends' feeds, not a new inbox. Meeting friends where they already are removed the biggest adoption risk: asking five people to install an app.
- Defer purchasing and retailer deals. Tempting for the business case, but they'd have doubled scope and diluted the single job the MVP had to prove.
Risks I planned for
Privacy of the social graph
Only the user's top-20 interactions are read, and only the five they choose are ever contacted. Nothing posts without an explicit send.
Slow or no responses
User-set response time plus a visible countdown so a silent friend never blocks a decision; partial results shown at timeout.
Social pressure & herd effect
Friends vote independently and don't see others' answers until the asker does, protecting honest opinions.
Eight-week clock
Weekly stakeholder demos and a written decision log kept scope honest and made every trade-off visible early.
Comforting, engaging, minimum fuss.
Wireframe exploration
Key screens were explored in low fidelity to find the most usable, scalable structure. Testing and refinement shaped a unified layout that works across user types and use cases.





Branding guidelines
Lo-fi sketches locked hierarchy, design goals and component choices before the visual layer was applied against the Opatter style guide.

Screens & interactions.
Splash and log in


Selecting your dearest friends


Making a purchase decision


Requesting an opinion — the rubber-band gesture
An alternative, one-thumb way to ask: drag the image down and it stretches like a rubber band, setting the response time as it goes. Release, and it flies to your friends.





What friends see

Instant decisions


Did it make deciding easier?
The MVP prototype was tested in two rounds with participants matching the persona, in simulated in-store scenarios: choose a gift, then choose something for yourself. Alongside task success I measured decision time, self-reported confidence, and a System Usability Scale score.
Measurement framework I set
Before testing I defined what "working" meant, so results could inform a go / no-go rather than a debate. One north-star, three supporting signals, tracked round over round.
| Level | Metric | Why it matters | Round 1 → 2 |
|---|---|---|---|
| North star | Decisions completed with a friends' verdict | Proves the core loop, not just the UI | 6/10 → 9/10 |
| Speed | Snap-to-sent time | Must fit an in-aisle moment | 2m10s → <90s |
| Confidence | Self-rated confidence after verdict (1–7) | The value proposition is confidence | +1.4 → +2.1 |
| Usability | System Usability Scale | Benchmarkable against industry (68) | 63 → 74 |
What changed between rounds
- Renamed the verdict buttons from generic thumbs to "Awesome / Meh" after testers found them warmer and faster to tap.
- Capped the circle at five after larger groups produced split, slower results.
- Added a visible countdown so the asker knows when to stop waiting and decide.
- Kept the slider as an accessible fallback for users who found the drag gesture hard.
Let the users inform the design.
By grounding every decision in user needs and best practice, the idea-pitch app delivered an onboarding and decision flow that built trust with new users and communicated its benefits clearly. The iterative loop of research, design and testing produced an interface people found intuitive on first use.
Ask the right questions
Directed but open questions validate hypotheses. Without them, interviews drift into vague answers that never test your assumptions.
Testing never ends
The testing phase doesn't finish when user testing does. Balancing user needs against time and resources, keep testing — it always leads to a better product.
You are not your designs
If a design doesn't improve the experience or meet a goal, it's decoration. Stay open in testing; the hard moments are where the pivotal changes come from.
Impact beyond the product
What I'd do differently as a manager
- Bring engineering into research sessions from week one, not only feasibility checks — shared empathy shortens every later conversation.
- Define the north-star metric before recruiting participants, so round one is a baseline rather than a rehearsal.
- Recruit a slightly older cohort earlier; the 27–42 segment surfaced different trust concerns that arrived late in the sprint.
- Plan the post-MVP roadmap alongside the pitch, so a "yes" from leadership has a next step ready.
Vibe Designing: a framework that turns evidence into working prototypes.
In close partnership with Intuit UX teams, I mentored and upskilled my team on Vibe Designing — a repeatable, AI-native framework that speeds up UX delivery and provides dev-ready code for engineering. Together, we synthesized PRDs, research, and stakeholder inputs via NotebookLM, generated actionable requirements and working prototypes using Claude, and systemized validated UI through Figma AI. By shifting from static mockups to fully traceable, runnable prototypes, we transformed how design collaborates with engineering. Here is a step-by-step look at the framework and its application on the TurboTax Pricing Source of Truth.
The handoff was the bottleneck.
Across IES, PPM and Compliance, my team was producing good Figma work — and losing a lot of it in translation. Research synthesis took days. Static mockups left engineers guessing at states, edge cases and logic. Every "why did we design it this way?" question sent someone back through a Zoom recording. Rework arrived late, when it was most expensive.
The opportunity wasn't a new tool; it was a new operating model — what the team now calls Vibe Designing. Generative AI had become good enough to transcribe, synthesise, draft and even build. If we re-sequenced the work around those capabilities — with the designer firmly in the loop as editor and decision-maker — we could hand engineering something far more useful than a picture: a working prototype whose every element traced back to customer evidence.
Typical time from a round of customer sessions to a shareable synthesis, before the change
Of engineering clarification questions were about states and logic a mockup couldn't show
Vendors, two toolchains, one product — consistency depended on individuals, not process
What I owned as the lead.
Changing how eight people across two vendors work is a leadership problem before it's a tooling one. I owned the operating model, the guardrails that made it safe inside a regulated financial-software company, and the upskilling that made it stick.
Framework & operating model
Authored the Vibe Designing framework — evidence → insight → requirement → prototype → system — and the gates between stages.
Governance & data safety
Agreed with security and legal what can enter which tool; set the PII-redaction and human-review rules every artefact follows.
Stakeholder alignment
Won engineering and PM buy-in by piloting on one feature, measuring, and letting the numbers make the case.
Upskilling & coaching
Ran weekly prompt-and-prototype clinics; paired every designer through their first vibe-coded prototype.
How the team runs now
- Synthesis within 24 hours of any customer session, reviewed live with PM.
- Weekly Vibe Designing clinic: share prompts and prototypes that worked, retire what didn't.
- Prototype review with engineering before any Figma polish.
- Monthly tooling retro with security on what data touched which tool.
- Evidence first: no requirement without a linked transcript excerpt.
- Human in the loop: AI drafts, a designer decides, a peer reviews.
- Prototype is the spec: states, logic and copy live in runnable code plus a short decision log.
- Figma is for the system: components and tokens, not pixel-perfect one-offs.
Vibe Designing, step by step.
Vibe Designing is the name my team gives to designing with generative AI as a disciplined, evidence-first practice — not prompting for pretty screens. Every run starts from the same inputs, moves through the same eight steps, and ends with a working prototype, a decision log and system components. The framework is tool-agnostic in principle; today it runs on Zoom, NotebookLM, Claude and the Figma AI suite.
Inputs the framework consumes
The pipeline below is the spine of the framework — every input on the cards beneath it enters at the first node and leaves engineering as a runnable prototype with its evidence attached.
The eight steps
Frame the problem & the evidence set
Agree the outcome with PM and PD, then assemble every relevant input — PRDs, mocks, meeting recordings, brainstorm boards, tickets — into one evidence set. Decide the north-star metric before anything is generated.
Capture & transcribe conversations
Record Zoom sessions with PD, PM, stakeholders and customers (with consent). Generate transcripts, redact names and account data, and tag each recording with its purpose so quotes stay attributable.
Synthesise in NotebookLM
Load the evidence set as sources. Interrogate it: pain points, jobs, constraints, contradictions between documents and conversations. Produce a cited insight brief, a current-state journey and open questions — every claim linked to a source.
Specify with Claude
From the brief, generate user stories, acceptance criteria, states and edge cases. The designer edits line by line; nothing passes without a linked citation. Open questions become explicit assumptions in the spec.
Vibe-design the prototype
Prompt Claude to build a working prototype in code on the design-system tokens — real states, validation, data shapes. Iterate in minutes, not sprints. The prototype is the spec: logic lives in it, with a short decision log alongside.
Review with PD, PM & stakeholders
Walk the working prototype in a Zoom review before any visual polish. Capture decisions and objections into the log; loop back to steps 3–5 as needed. Peers brainstorm alternatives against the same evidence.
Validate with users
Test the prototype with target users on real tasks. Measure task success, time and confidence; re-run the generation loop on what fails. Findings go back into NotebookLM as new sources.
Systemise in the Figma AI suite
Convert the validated prototype into components, variants and documentation mapped to the Intuit design system. Hand engineering the prototype, the spec, the decision log and the system components together.
Where each tool earns its place
| Stage | Tool | What it does for us | Guardrail |
|---|---|---|---|
| Capture | Zoom recordings & transcription | Every session becomes searchable text the same day; verbatim quotes survive into the spec | Consent captured; names and account data redacted before anything leaves Zoom |
| Synthesise | NotebookLM | Grounded Q&A across dozens of transcripts; insight briefs, personas and FAQ with source citations | Only redacted sources; outputs checked against citations by a second designer |
| Specify & build | Claude | User stories, acceptance criteria, edge cases, then a working prototype in code on our tokens | Approved enterprise workspace; no customer data in prompts; designer owns every decision |
| Systemise | Figma AI suite | Prototype to components fast; variants, auto-layout and docs generated and curated | Everything mapped to the Intuit design system before publishing |
How to do Vibe Designing well.
Speed is the easy part of designing with AI. The hard part is staying trustworthy, consistent and honest while moving fast. These are the practices I hold my team to — and the ones I ask engineering and product partners to hold us to. They are written as rules on purpose: a framework is only as good as the discipline around it.
1 · Set guardrails before you generate anything
Guardrails are what turn an experiment into a practice an enterprise can trust. In a regulated financial-software company they are also the licence to operate: security, legal and compliance said yes to Vibe Designing because the rules were written down first.
Data boundaries
Only approved enterprise workspaces. Customer names, account data and financial figures are redacted before a transcript or document enters any tool. Nothing confidential in a prompt — ever.
Human in the loop
AI drafts; a designer decides; a peer reviews. No generated requirement, copy or prototype reaches a stakeholder without a named reviewer and a reason it was accepted.
Evidence or it doesn't ship
Every insight carries a citation to a source; every requirement links to the insight. If a claim can't be traced, it's treated as a hypothesis and labelled as one.
Label what the AI made
Prototypes, drafts and syntheses are marked as AI-assisted and as prototype-grade. Engineering owns production code; the prototype is the spec, not the shipment.
Why guardrails come first
Three reasons, in the order leadership cares about them. Trust: one leaked customer detail or one confidently wrong requirement would end the programme and set AI-assisted design back years. Speed: teams move fastest inside clear boundaries — nobody pauses to ask "am I allowed to?" when the answer is written down. Scale: rules let eight people across two vendors behave as one team, and let a new joiner be productive in a week.
The guardrails charter we signed with security, legal and engineering
| Guardrail | What it means in practice | How it's enforced | Owner |
|---|---|---|---|
| Approved tools only | Enterprise Claude, NotebookLM and Figma workspaces with no training on our data; consumer tiers are off-limits | SSO-only access; tool list reviewed quarterly | Security |
| Redact at source | Names, emails, account IDs, pricing figures and anything contractual are removed before transcripts or docs leave Zoom or Drive | Redaction checklist on every source; spot-checks in the monthly review | UX lead |
| No confidential data in prompts | Prompts describe structure and intent; real customer or financial data is never pasted in | Prompt templates with placeholder fields; peer review of shared prompts | Every designer |
| Human sign-off | A named designer approves every generated requirement, copy block and prototype before it reaches a stakeholder | Reviewer field in the decision log; nothing shared without it | Designer + peer |
| Cite or label | Every insight points to a source; anything untraceable is labelled a hypothesis | NotebookLM citations checked; "hypothesis" tag in the spec | Designer + peer |
| Prototype ≠ production | Prototypes are labelled AI-assisted and prototype-grade; engineering builds production from the spec | Banner in every prototype; handoff checklist | UX lead + Eng |
| Bias & harm check | Generated personas, copy and defaults are reviewed for stereotyping, exclusion and dark patterns | Inclusive-design checklist in critique | Critique group |
| Log exceptions | Any near-miss or breach is recorded and discussed, not hidden | Monthly tooling retro with security | UX lead |
2 · Connect to the Intuit design system and guidelines
A vibe-designed prototype that ignores the design system is just a faster way to create rework. We give the models the system up front so what they generate is already ours.
Tokens and components as context
Design tokens, component specs and content guidelines are loaded as sources in NotebookLM and attached to Claude as project knowledge, so colour, type, spacing, states and tone are inherited — not reinvented.
Prompt from the system
Prompts reference named components ("use the IDS data table with inline edit") rather than describing visuals. If a pattern doesn't exist, that becomes a design-system conversation, not an improvised one.
Accessibility and content standards
WCAG 2.1 AA, Intuit's content voice and localisation rules are part of the acceptance criteria generated in step 04 — so they're checked in review, not discovered in audit.
Give back to the system
Anything validated and new goes through the Figma AI suite into the system backlog with its evidence, so the next team starts from a component, not a prototype.
How the system is wired into the tools
- Design tokens as code. Colour, type, spacing, radius and elevation tokens are provided to Claude as a tokens file, so generated prototypes compile against the system rather than invented values.
- Component documentation as context. Component specs, states and usage rules from the Intuit design system are loaded as NotebookLM sources and attached as project knowledge, so the model can answer "which component?" with our answer.
- Reusable skills. Common tasks — "data table with inline edit", "approval status pattern", "empty and error states" — are captured as reusable prompt skills the whole team shares, versioned alongside the design system.
- Content and accessibility guidelines in the acceptance criteria. Voice, terminology, reading level, WCAG 2.1 AA and localisation rules are generated into every story, so review checks them explicitly.
- Figma AI suite as the bridge. Validated prototypes are converted into system components, variants and documentation, keeping Figma as the source of truth for the system.
A prompt pattern that respects the system
- Outcome & evidence — the user goal, the north-star metric, and the cited insights this screen must serve.
- System constraints — "use the Intuit design system components and tokens provided; do not invent colours, type or spacing; name the component you use for each element."
- Content & accessibility — voice rules, terminology list, WCAG 2.1 AA, keyboard and screen-reader behaviour, localisation.
- States & logic — empty, loading, error, permission and edge cases from the acceptance criteria; data shapes to use.
- Uncertainty — "list assumptions, open questions and your confidence per requirement; flag anything not covered by the evidence."
- Output — a runnable prototype plus a short decision log in plain English.
When the system doesn't have the pattern
Don't improvise in the prototype. Raise it as a design-system proposal with the evidence attached, prototype it as a clearly labelled candidate, and let the system team decide. The TurboTax approval-state pattern (Proposed → Finance Review → Approved → Applied) went through exactly this route and is now a shared component rather than a one-off.
3 · Keep the skills that make the judgement
Vibe Designing raises the bar on craft rather than lowering it. The model can draft; it cannot know what matters. These are the skills I hire for, coach and protect in the team.
What "good" looks like, by level
| Skill | Designer | Senior | Lead |
|---|---|---|---|
| Research & synthesis | Runs a session; checks citations in a brief | Designs the study; spots unsupported claims | Sets the research plan and evidence standard |
| Prompt & context design | Uses shared skills and templates | Writes new skills; supplies the right sources | Curates the skills library; sets guardrails |
| Reading & shaping code | Reviews states and copy in a prototype | Adjusts logic and tokens; pairs with engineers | Judges feasibility and handoff quality |
| Systems & design-system fluency | Uses the right component | Proposes patterns through the system | Owns the relationship with the system team |
| Content design | Edits generated copy to voice | Writes terminology and acceptance criteria | Sets content standards for the team |
| Facilitation & critique | Presents evidence clearly | Runs reviews where evidence decides | Aligns PD, PM and stakeholders on trade-offs |
How we grow them
- Every designer is paired through their first vibe-designed prototype by someone who has shipped one.
- Weekly clinic: one prompt that worked, one that failed, and why — recorded into the skills library.
- Critique is unchanged in cadence and standard; AI-assisted work is reviewed as work, not as a demo.
- Engineers join clinics monthly so code literacy and feasibility judgement grow together.
- Skills show up in goals and growth conversations, so adoption is coached, not mandated.
4 · Ask when in doubt
The failure mode of AI-assisted work is quiet confidence. The antidote is a culture where asking is faster than guessing, and where the tools themselves are made to surface uncertainty.
Interrogate NotebookLM before asserting: "What contradicts this?", "Which sources say this?", "What did Finance actually say?" If the brief can't answer, go back to the people.
Prompts require Claude to list assumptions, open questions and confidence per requirement. Low-confidence items are flagged in the spec, never silently filled in.
Open questions have an owner and a date. On TurboTax, "where does pricing live in PPM?" and "how many simultaneous editors?" were raised in week one, not discovered in build.
Weekly Vibe Designing clinic and a "stuck?" channel with a same-day answer norm. Nobody loses a day to a prompt that won't behave or a pattern they're unsure of.
A simple rule for when to stop and ask
- a requirement has no citation and you'd be inventing the "why"
- the model's output contradicts a source, or two sources contradict each other
- a pattern isn't in the design system
- a decision touches money, permissions, compliance or data
- you've iterated a prompt three times without progress
- you're about to paste anything that might be confidential
- the evidence is cited and consistent
- the component and content rules are clear
- the change is reversible and labelled prototype-grade
- the assumption is written down and flagged in the spec
Asking is cheap; the cost is a message. Guessing is expensive; the cost is a sprint. The team's norm — ask within the hour rather than guess for a day — exists because AI makes guessing feel safer than it is.
The team's working rules, in one place
Start from the outcome and the evidence set, never from a prompt.
Redact first. Nothing confidential enters any tool.
Cite or label it a hypothesis.
Prompt from the design system; propose new patterns through it.
Review the working prototype before any polish.
Make assumptions and confidence visible in every spec.
Ask within the hour rather than guess for a day.
Hand off prototype + spec + decision log + components together.
Craft gates don't move: critique, accessibility and content review stay mandatory.
What engineering actually receives.
A Figma file and a meeting
- Happy-path screens; states and errors implied or missing.
- Logic described in comments and Slack threads.
- "Why?" answered from memory or by re-watching recordings.
- Rework discovered in sprint review.
- Prototype fidelity limited to click-through.
A runnable prototype and a decision log
- All states, validation and edge cases working in code.
- Acceptance criteria generated from evidence and approved by design.
- Every element traces to a customer quote in the brief.
- Rework caught in prototype review, before the sprint.
- Tested with real data shapes and real latency.
Worked example: TurboTax Pricing Source of Truth
Phase 1 of a programme to replace TurboTax's spreadsheet-based pricing workflow with a single, trusted pane of glass in PPM, for Consumer Group Pricing Ops, CG Strategy and Finance. The evidence set: the UX & Design Requirements deck, the existing Google Sheet workflow and PPM screens, Zoom working sessions with PD, PM, Pricing Ops and Finance partners, two peer brainstorms, and a short set of operator interviews.
- Fragmented views — pricing spread across baseline, desktop, BizTax and full-service tabs with no unified product view.
- Manual, repetitive entry — every change means inserting columns and copy-pasting unchanged values.
- No structured approval — Finance validated via sheet comments: easy to miss, impossible to enforce.
- Hard to share — leaders got a spreadsheet link, not a governed view.
- One pane, one truth — PPM replaces the sheet as where pricing lives.
- Fewer steps to update — editing a price should be easy, not a ten-step form.
- Controlled, accountable access — authorised editors only, finance validation built in.
- Personalised, not one-size — filter and save the view that matters to your role.
How Vibe Designing ran on this project
The vibe-designed output
A one-minute screen recording of the prototype being vibe-designed in Claude — requirements in, working PPM dashboard out. Watch the unified price view, the four-state finance approval flow, inline editing with record locking and governed sharing take shape.
What good looked like at launch
Discovery to a validated, engineering-ready prototype took nine working days, against roughly four weeks on the previous Figma-first model for comparable PPM work.
Making it safe and stick.
Decisions & trade-offs I made
- Prototype before Figma, not after. Reversing the order felt wrong to the team at first; it is where most of the turnaround gain came from.
- Pilot on one feature, then scale. I resisted a big-bang rollout. One PPM feature, measured honestly, made the case to PM and engineering better than any deck.
- Evidence links are mandatory. A requirement without a transcript citation is sent back. Slower for a week; far fewer "who asked for this?" debates since.
- Keep Figma for the system. We stopped producing pixel-perfect one-off screens; the AI suite now builds components from validated prototypes instead.
- Approved tools only. We waited for enterprise workspaces rather than using consumer tiers, even though it cost us weeks at the start.
Risks I planned for
Customer data leaving the boundary
Redaction before transcripts move; enterprise workspaces with no training on our data; monthly review with security.
Hallucinated insights or requirements
NotebookLM citations checked by a second designer; Claude outputs reviewed line by line before they become stories.
Prototype mistaken for production code
Prototypes are labelled and scoped; engineering owns the production build; the prototype is the spec, not the shipment.
Craft erosion and uneven adoption
Critique stayed weekly; clinics paired fast adopters with sceptics; design quality gates unchanged.
The impact on delivery.
I agreed a scorecard with PM and engineering before the pilot so the process was judged on delivery outcomes, not on how novel the tools felt. Measured across IES, PPM and Compliance work over two quarters.
How I measured it
| Level | Metric | Why it matters | Movement |
|---|---|---|---|
| North star | Features engineering-ready per quarter | Throughput of validated, buildable work | +30% turnaround |
| Quality | Design-related clarification tickets & rework | Did the handoff actually get clearer? | −15% rework cycle time |
| Speed | Session-to-synthesis; brief-to-prototype | Where the hours went before | 3 days → 4 hours; weeks → days |
| Adoption | Designers shipping vibe-coded prototypes | The process only works if everyone uses it | 8 of 8 within one quarter |
| Safety | Data-handling exceptions logged | Licence to keep operating | Zero |
Impact beyond the metrics
What leading an AI-native team taught me.
Sequence is the strategy
The tools mattered less than the order. Evidence → spec → prototype → system is what removed the rework, not any single model.
Guardrails buy you speed
Agreeing data rules with security up front turned a potential blocker into the reason leadership let us move fast.
Coaching beats mandating
Pairing each designer through their first prototype did more for adoption than any policy. Sceptics became the best teachers.
What I'd do differently as a manager
- Baseline the clarification-ticket count a quarter earlier, so the quality story started from clean data.
- Bring one engineer into the prompt clinics from day one; they became the strongest advocates once they joined.
- Write the data-handling playbook before the pilot rather than during it — it would have saved two weeks of waiting.
- Plan the system-level Figma work as a parallel track, not a final step, to avoid a documentation backlog after each prototype.