Draft · work in progress

Design Story · Flagship · AI experience design

LEDGER

Designing trust into an AI policy assistant used by 250,000 people. As sole owner of the UI and UX, I turned a promising language model into something policy knowledge-workers could actually rely on — where every answer shows its sources, holds up to scrutiny, and survives a skeptical room of policy, legal, and executive stakeholders.

95%Reduction in lookup time
17Usability tests
Placeholder plate — redacted flows & wireframes to be added

The brief

A U.S. Department of Homeland Security team set out to put an AI assistant in front of hundreds of thousands of employees to help them navigate dense, high-stakes policy. The technology worked. The question was whether people would trust it enough to use it — and whether the experience could earn the confidence of the policy, legal, and executive stakeholders who had to stand behind every answer it gave.

The problem, reframed

Usability testing with policy knowledge-workers surfaced the real requirement early: source transparency and response precision were non-negotiable. An answer without a traceable source wasn’t just less useful — it was unusable. That single finding reframed the work from “design a chat UI” to “design a system people can verify.”

What I decided

Findings drove the interaction model toward visible, checkable sourcing; motivated architecture changes to support multi-turn conversation, so users could refine and pin down an answer; and shaped a CRIT prompting guide — Context, Role, Interview, Task — built to teach policy knowledge-workers to get better answers in less time. That same guide doubled as our UAT test scenarios: the prompts that modeled good use were the ones we validated against. I also resolved navigation and discoverability issues before launch.

The outcome

LEDGER shipped to roughly 250,000 users with buy-in built across a skeptical coalition of policy, legal, and senior executives. In testing, the experience cut the time to find and verify a policy answer by about 95% against the old multi-document slog — the product of 17 rounds of usability testing that hardened the interaction model before launch.