Keith Cox, PhDBehavioral Researcher & AI Practitioner

Portfolio

Design Prototypes

End-to-end system design case studies, documented through the architecture itself.

Pharmacy Voice Agent

2025

Figma + Voiceflow + Airtable

Complete multi-agent architecture with identity-gated data disclosure and deliberate human escalation, documented end to end.

Designed a multi-agent voice bot for a fictional pharmacy: wireframed in Figma, built it in Voiceflow, and backed with a synthetic retail pharmacy database in Airtable. A routing agent triages every call into FAQ, patient-specific, or clinical questions, and the architecture is built around a strict rule: no prescription data leaves the bot until identity is verified, and no clinical judgment is ever made by the bot at all. Built and tested in Voiceflow; the trial's free-tier credits have since run out, so this entry documents the design through the full agent flow diagram rather than a live link.

  • Identity gate: patient-specific requests require name + date of birth to be verified against Airtable records via live API call before any prescription information is disclosed; when callers can't be verified successfully, the agent first explains that private patient information cannot be disclosed to unverified callers, then offers to transfer the call to a human.
  • Live lookup: verified callers' queries and requests are resolved using data retrieved from the Airtable pharmacy database via a live API call.
  • Status classification: a classification agent sorts each inquiry — ready for pickup, awaiting prescriber approval, expired, insurance delay/prior authorization, or other — and routes to the matching response.
  • Human escalation: clinical or medication questions are deliberately kept out of the bot's scope and escalated to a human pharmacist rather than answered.

Figma wireframe of the full agent flow.

Browse the synthetic pharmacy database

Web Apps / Technical Projects

Real, live software, designed, built, and deployed end to end.

Conference Scheduler

2026
icca2026.vercel.app

Scoped, designed, and built on the fly to address flaws with the official conference app. Achieved conference-wide adoption in under 24-hours.

Identified core failures in an international research conference's official app (no talk-level search or scheduling) and built a better alternative in a single afternoon with Claude Code: full-text talk search, per-talk scheduling, and a live "happening now" view keyed to the event's 30-minute talk grid. ICCA convenes once every four years, so the official app's developers had months if not years of lead time by comparison. The webapp spread through the conference like wildfire, such an obvious improvement over the official app that even conference organizers adopted it and recommended it to attendees.

Search comparison: My app (left) vs. the official ICCA 2026 conference app (right).

Some-R-Set Scorepad

2026
somerset-scorepad.vercel.app

Live in production: players join by code, spectate games in real time, and every game persists to shared history.

Real-time multiplayer scorekeeping app for a partnership card game, built with Claude Code and shipped as a single dependency-free HTML file — vanilla JavaScript, no framework or build step — that installs as an offline-capable web app. It automates the bid-and-set arithmetic players routinely get wrong: hand-by-hand bid and take entry, automatic set penalties and shoot-the-moon detection, and dealer rotation. A two-phase entry flow lets one player log a bid while the whole table spectates the hand live, then any device records the result. Firebase drives live cross-device sync — players join a game or tournament with a six-character code — alongside opt-in cloud backup with device linking and privacy-preserving stats sharing between players. On top of the core scorepad sit three tournament formats (single elimination, double elimination, and round robin) and a Personal Stats engine that computes leaderboards, win streaks, moon-conversion rates, and best-partner and nemesis records retroactively from game history.

Scorepad live game board mid-hand, showing team scores, the bid/take rotation, and recent deal history
Scorepad single-elimination tournament bracket with the next matchup and completed round results
Scorepad game history log with expandable hand-by-hand results and a delete-game option
Scorepad player stats view showing win-loss record, championships, streaks, and moon shot success rate

Self-Hosted LLM Lab & RAG Research Assistant

2025–Present

Runs a local LM Studio environment for systematic prompt-engineering experimentation. Built a retrieval-augmented generation (RAG) personal research assistant over my own academic literature corpus that surfaces relevant studies based on working questions, using few-shot prompting informed by conversation-analytic method to specify not just desired outputs but the underlying interactional move an agent should recognize or perform.

Conversation Design Philosophy

Grounding conversation design in the empirical study of real talk.

My conversation design philosophy is a reflection of my training as a social scientist in general and my specialization as a conversation analyst, specifically. Conversation analysis is the discipline of specifying what "good" looks like in an interaction: not just what people are saying, but what they are doing, and exactly how they are doing it. That is precisely the unsolved problem at the center of AI experience work — defining, evaluating, and improving how an agent should behave, turn by turn.

My work brings decades of conversation-analytic research on naturally occurring talk to bear on this problem: I leverage what we know about how humans interact with each other to identify specific failure modes in conversational interfaces, recommend concrete engineering solutions, and build evaluation protocols that are rigorous enough to measure the effects of both.

Designing for Trust: What Makes an AI Conversation Feel Trustworthy

1. What makes an AI interaction feel trustworthy

Trust has to do with predictability, reliability, and transparency. Humans trust each other when their conduct is relevant, reasonable, and accountable. Three levers do most of the work: fitted responses (recipient design), graceful repair (especially user-initiated repair), and a calibrated epistemic stance (not only keeping track of who knows what at any given moment but modifying the content and design of turns accordingly in real time).

2. Where AI conversations break down

"Mechanical listening" instead of action-first listening (i.e., focusing on what the user is saying rather than on what the user is doing); treating resistance as confusion; overly rigid sequential expectations; weak repair initiation; losing context across turns.

3. Precision prompt engineering

A conversation-analytic vocabulary produces better few-shot examples, because it specifies the interactional move an agent should recognize, not just the surface text.

Conversation analysis is the differentiator here, grounding conversation-design decisions in decades of empirical literature on how real talk actually works, not just intuition about what sounds natural.

Publications

Peer-reviewed research on trust, resistance, and breakdown in high-stakes conversations.

  • Cox, K. (2025). When Good News Falls Flat: Complications in the Delivery and Reception of Good News in Pediatric Neurology.Social Psychology Quarterly 88(1): 45–65. DOI: 10.1177/01902725241253258

    How well-intentioned clinical communication fails to land, and the conversational moves that repair it.

  • Timmermans, S., Stivers, T., Cox, K., & McArthur, A. (2024). Patients in Pain: How Treatment Plan Formulations Shape Patient Response.Communication and Medicine 19(2): 137–151. DOI: 10.1558/cam.22881

    How small changes in phrasing produce measurably different responses: the empirical case for prompt-level precision.

  • Cox, K. (2023). Invoking Uncertainty: Parents' Accounts for Intrusions on Medical Authority in Pediatric Neurology.Journal of Health and Social Behavior 64(4): 537–554. DOI: 10.1177/00221465231194052

    How people resist expert authority in high-stakes settings: patterns any trustworthy AI system must recognize and handle.

  • Tarn, D. M., Barrientos, M., Pletcher, M. J., Cox, K., Turner, J., Fernandez, A., & Schwartz, J. B. (2021). Perceptions of Patients with Primary Nonadherence to Statin Medications.Journal of the American Board of Family Medicine 34(1): 123–131. DOI: 10.3122/jabfm.2021.01.200262

    Why people abandon prescribed plans: the human analog of drop-off and task-completion failure in automated workflows.

In progress

Interactional Risk in First-Time Conversations Between Strangers

UCLA collaboration, manuscript in preparation: conversation-analytic study of 140 speakers across 70 dyads in first-time conversations, examining how speakers calibrate interactional risk when making evaluative claims without shared relational history. Directly analogous to agent–patient first contact.