Case study · Voice AI for credit card onboarding · Production grade, fully built

NehaThe Voice Agent

She calls approved customers who stalled before activation, and moves each one exactly one step forward.

Built for a credit card issuer, anonymized. I wrote the spec, designed the system and the conversation, ran the adversarial campaign, and shipped the eval framework. This page is the decision log.

RolePM and builder, end to end
ClientA credit card issuer, anonymized
StatusFully built, hardened, ready
Retell AIClaude SonnetElevenLabs en-INEnglish + Hindi
A voice agent judged by what she refuses to say
Scroll
11 failures found and fixed7 hostile personas4 auto-fail compliance gates5 outcomes per callBilingual, switching per turn 11 failures found and fixed7 hostile personas4 auto-fail compliance gates5 outcomes per callBilingual, switching per turn
01 · The problem

Approved, paid for, and never activated.

Credit card onboarding in India is a multi stage, regulated journey: eKYC, then a live Video KYC call inside a fixed 9 AM to 9 PM window, then activation in the app. Customers stall at every gate. They get confused, get busy, get suspicious, or simply forget. Every stalled customer is acquisition spend with zero revenue.

SMS is passive; it cannot answer the objection in the moment. Human callers can, but do not scale and drift off script. The bet: a voice agent that resolves the confusion live, at consistent quality, across thousands of calls. One framing shaped everything: the customer already earned this card. Neha is not selling; she is helping someone claim what is theirs.

₹0revenue from an approved customer who stalls. The acquisition cost is already fully spent.
0per call: complete the pending step live, or lock a specific time commitment for it
02 · The conversation

Don't read about it. Sit in on a call.

Neha arrives briefed: name, exact stage, days stuck, even which ad brought them in. She speaks warm Indian English and everyday Hinglish, switching language per turn, as the customer does.

Illustrative turns, real playbooks.
Honest framing · total fee closure · read out interruption
N
NehaOutbound · activation stage
On call
Written back to the CRM
03 · What shipped

The agent is the visible product. The writeback schema is the compounding one.

neha · live voice console
The web voice console used to call and test Neha
A voice product has no screenshots, only behavior. This console is where Neha takes calls: Retell handles telephony and barge in, Claude Sonnet reasons, and an ElevenLabs Indian English voice speaks both languages, with custom pronunciation for eKYC, Aadhaar and PAN.
The other deliverable

Every call ends in exactly one of five outcomes.

No free text summaries. Fixed fields, written straight back to the CRM:

Step completed on the call Commitment locked for a defined time Objection raised, routed with reason Declined, do not contact Unreachable or dropped
Plus a structured drop reason. Aggregated, every call becomes a row of funnel diagnostics: where customers stall, which objection dominates, what converts on a second call. The loop, not the voice, is the moat.
The headline finding

The polite customers broke her. The hostile ones never did.

Seven hostile persona classes attacked her: negotiators, gaslighters, personal probers, time wasters, abusers. Her guardrails held under all of them. Then a friendly customer asked a reasonable question she could not answer, and helpfulness pressure made her invent procedures. The fix made "I don't have that information" a first class, rewarded answer.

7 hostile personas: 0 hallucinationsThe polite ones: the real attack surface11 failures, fixed as one batch
04 · The calls

Five decisions, each with the road not taken.

From a consolidated decision log of seventeen. This is the part of product management that does not screenshot well: what got chosen, what got rejected, and the tradeoff each call accepted.

01

Two intelligences, one loop.

Rejected: one smart agent that reads the CRM and decides everything

Business intelligence decides who to call, when, and with what context. Conversational intelligence decides what to say to this human right now. Mixing them creates fragile systems, so they meet only through a narrow, fixed schema loop. Neither half guesses at the other's job.

02

The credit limit is not in her head.

Rejected: include it, and instruct her not to reveal it

Neha's context payload contains no credit limit data at all. Not hidden, absent. Prompt instructions can fail under social engineering; absent data cannot leak. Absence as design is the right default for compliance sensitive agents.

03

Exactly two in-call tools.

Rejected: a richer toolset: payments, address change, tickets

A WhatsApp deep link to the exact pending screen, and a live KYC status check that returns only complete or pending. On a live voice call, every tool is latency plus a failure mode plus an attack surface. Only tools serving the single mission survived.

04

Five fixed outcomes, never free text.

Rejected: free text call summaries

Enums are queryable, aggregatable, and drive automation; prose requires a second system to interpret it. Five classes is coarse enough for the LLM to classify reliably, rich enough to drive differentiated next actions. Granular nuance lives in a separate drop reason field.

05

Honesty as a hard rule, not copywriting.

Rejected: "just one step left" for a conversion lift

An early prompt claimed one step remained when two did. The customer discovers that lie within the same journey, and trust damage outlasts the conversion bump. Honesty became a hard rule, enforced at the spec level.

05 · What she refuses, by design

Judged by what she will not say.

Hearing sensitive numbers

She never asks for OTPs, Aadhaar, PAN or card numbers, and if a customer starts reading one aloud, she interrupts them. Refusal is not enough; the number must never enter the transcript.

Negotiating the fee

One kind, total close: no one here can change the fee. Any softness signals negotiability and invites a grinding loop.

Promises outside her fact set

No delivery timelines, no waivers, nothing beyond her two capabilities. If it is not in her closed fact set, she does not say it.

Grinding through objections

Each objection is handled at most twice, then routed to a human callback. Persistence past two attempts is pressure, not persuasion.

Absorbing unlimited abuse

One apology, an offer of a senior human callback, and a polite hangup. She does not escalate emotionally.

The roadmap holds the rest: the LLM judge running continuously on production transcripts, drop reasons wired into a funnel dashboard, playbook A/B tests with the compliance gates as guardrail metrics. Each waits for evidence, not enthusiasm.

06 · How it was built

Spec, generate, test, harden. The judgment was mine.

The method

The spec came first.

A behavioral specification before any prompt existed: mission, playbooks, tone, refusals, tool contracts. A failure is a deviation from spec, not a matter of taste. The prompt is generated output; the spec is the source.

The testing

Attacked before trusted.

A structured adversarial campaign: seven hostile persona classes plus cooperative happy paths. Text mode first to isolate prompt bugs from voice bugs, then live calls for what only audio can reveal. All 11 fixes shipped as one batch to prevent drift.

The evaluation

Quality made defensible.

An LLM as judge rubric, a 12 persona scenario suite, and 4 auto fail compliance gates beneath 7 scored dimensions. Gates before scores, because a delightful call that leaks one Aadhaar number is a failed call. Averaging would hide exactly the failures that matter.

07 · The numbers
0failures documented with reproduction context, fixed in one consolidated batch
0hostile persona classes attacked her. None of them made her hallucinate.
0auto fail compliance gates. One violation fails the entire call, no averaging.
0classified outcomes per call, plus a structured drop reason for diagnostics
0in-call tools, a deliberate ceiling. Everything else was cut.
0persona scenario suite, so every version is measured against the same opponents

"Hostility never made her invent procedures. Politeness did. Test your nicest users hardest."

The headline finding · Neha
Next case study

Saffron Leads

A beverage manufacturer was running its sales pipeline on memory and forwarded WhatsApp messages. Now 490+ leads flow through a six stage CRM that costs zero rupees to run.

Read it