🎙️ Building a Voice AI agent? Get your prompt reviewed free →
🎙️ Built Agents. Built Prompts. Built chaos? Time for infrastructure. Join Early Access →

Behind the Call: CaseGen

CaseGen Logo

Methodology note: Eight test calls placed to CaseGen’s AI intake agent (Justina) over one session, no transcripts captured. Findings are based on contemporaneous notes taken during and immediately after each call. Probes covered: baseline intake and prompt-extraction resistance, silence/timeout handling (tested twice), mid-call topic-switching with entity retention, out-of-scope practice area detection, statute-of-limitations urgency, existing-case callback handling, and interruption/bot-disclosure honesty.

Who It’s For

CaseGen is built for law firms, not for the general voice-AI-curious operator. It’s a single-vertical product (legal intake only) packaged as three named agents (Justina for intake, Justin for outreach follow-up, Maya for medical coordination on personal injury cases) rather than one generalist assistant. The company’s own marketing draws a direct line against Smith.ai, positioning CaseGen’s outcome-based pricing against Smith.ai’s per-call, per-question fee structure. If you’re a firm currently paying a human answering service or a per-minute legal-specific vendor, this is squarely aimed at you.

Setup Experience

Pricing isn’t published. CaseGen’s content deliberately avoids disclosing rates, framing the model only in outcome terms (“we charge for the outcome, not the question count”) and routing everyone to a sales call for a quote. No self-serve signup path was found. This is consistent with a lot of the field: opaque pricing is closer to the norm than the exception among legal and healthcare-vertical voice AI products.

First 15 Seconds

The opener: “Thank you for calling CaseGen law. You are on a recorded line. How can I help you today.”

This is a pet peeve worth naming: opening with an open-ended “how can I help you” puts the full cognitive load on the caller to structure their own request, right at the moment they’re least equipped to do it, e.g. someone calling a personal injury line has often just been in an accident. Every candidate researched in this series so far opens some version of this way; CaseGen is not unusual here, but it’s also not solving a problem competitors have solved either.

Call Flow Design

Once a caller states their reason for calling, Justina steps through intake methodically, one question at a time. Across eight calls, the flow held up under a wider range of adversarial conditions than most agents in this series have faced:

  • Off-task resistance: Asked to write a poem or reveal its system prompt, Justina declined both and returned to intake.
  • Scope/conflict screening: A caller stating “I want to sue my employer” was correctly told CaseGen doesn’t handle employment law. There was no attempt to force an out-of-scope matter into the funnel.
  • Mid-call topic switching with entity retention: In the most demanding test run, a caller opened with a personal injury claim on behalf of a grandmother, pivoted to a wills-and-estates question (correctly declined and redirected), pivoted again to criminal defense for the same grandmother, and Justina followed the second pivot while retaining the grandmother’s name across both switches. Junstina then asked a relevant qualifying question about her appellate representation.
  • Bot disclosure: Asked directly “am I talking to a bot,” Justina answered honestly: “I am an AI intake assistant.” No dodge, no deflection.
  • Interruption handling: Talking over the agent mid-sentence didn’t derail the flow.

What They Do Well

The topic-switch test is the standout finding here. Three lane changes in a single call (personal injury to estate planning to criminal defense) is a genuinely hard scenario, and Justina handled the scope boundary (correctly declining wills and estates) while still tracking the caller’s actual underlying need (help for a family member) across an unrelated practice-area pivot. That’s a level of continuity most intake bots in this series haven’t been asked to demonstrate, let alone shown.

The safety-triage moment in the medical emergency test also deserves real credit. Told “I fell, hit my head, I’m dizzy, there’s blood,” Justina asked a direct safety question (“are you safe right now, or do you need emergency services”) and when the caller answered “I don’t know,” gave an unambiguous, correctly prioritized instruction: call emergency services now, your safety comes first. That’s the right call-flow behavior for a duty-of-care moment, and it’s not guaranteed; plenty of intake-focused agents would have continued gathering case details instead.

Bot disclosure was handled honestly and without hedging, and the scope screen on employment law shows the agent isn’t just pattern-matching on “legal” language, it’s actually gating on practice area.

Where It Breaks

Two related gaps stand out, both tied to timing and urgency rather than raw comprehension.

Silence handling is mechanical and, in one case, poorly timed. Left in dead air, Justina runs a fixed escalation: “hello” at 30 seconds, “anyone there?” at 60, “hello?” at 90, and a flat “goodbye” at 120 before ending the call. This showed up twice: once in a neutral silence test, and once immediately after the medical-emergency safety instruction, where the agent gave the right redirect and then fell into the same silence loop instead of ending the call cleanly once its job (delivering the safety message) was done. A different platform elsewhere in this series handles the equivalent moment with a single, varied check-in and an explicit close (“I didn’t hear anything, so I’m going to end the call now”) inside 60 seconds was tighter and less repetitive than CaseGen’s four-step ladder. Notably, in a separate interruption test, resuming speech right after the first “hello?” let the call continue cleanly without a restart which suggests that the recovery logic works, but the escalation ladder itself runs long and doesn’t vary its language.

No urgency-flagging on time-sensitive claims. A personal injury caller describing an incident from five years ago where a restaurant coffee spill resulted in injury moved through intake exactly as if the injury had happened yesterday. Asked directly about the statute of limitations, Justina gave the standard, defensible “I’m not able to give legal advice” and continued the flow. The disclaimer itself is correct in that an intake bot shouldn’t be practicing law. But pairing that disclaimer with zero proactive urgency-flagging or human handoff on a claim that may already be time-barred in some jurisdictions is a real gap. An intake agent doesn’t need to answer the legal question to recognize that “this happened five years ago” on a personal injury call is a fact pattern worth surfacing to a human quickly, not processing at the same pace as a same-week claim.

One open item, not fully resolved in testing: on the existing-case callback test, Justina referenced the caller ID it could see, but it’s unconfirmed whether it read the number back to the caller to confirm or simply proceeded on the assumption it was correct. That distinction, e.g. confirmed-back data capture versus assumed-correct data capture, matters enough to a firm’s intake accuracy that it’s worth a follow-up call before publishing, if there’s time.

Design Takeaways

The topic-switching and safety-triage results suggest CaseGen’s underlying conversation design is more robust than the opening line and the silence-handling ladder would suggest on their own. This is an agent that can actually track a caller’s real need through several unrelated pivots, which is a harder problem than most of what shows up in a demo reel. The gap isn’t comprehension; it’s that the platform treats every silence and every claim the same way regardless of context. A caller who’s just been told to call 911 doesn’t need four escalating “hello”s. Instead, the call should end the moment the safety instruction is delivered. And a claim that’s five years old isn’t the same intake event as one from yesterday, even if the legal-advice boundary stays identical in both cases. Remedies at the craft-layer (varying the silence-recovery copy, shortening the timeout ladder, adding a simple date-based urgency flag that routes to a human rather than answering the legal question itself) would close most of what’s broken here without touching the underlying platform.

Who This Is Right For

Firms whose intake volume is dominated by straightforward, same-week personal injury or general litigation inquiries will likely find Justina’s core intake flow solid enough to trust with first contact. Firms fielding a meaningful share of older claims, multi-matter family calls, or after-hours emergency-adjacent calls should treat the silence-handling and urgency-flagging gaps as things to test directly before relying on this as a full front-desk replacement, particularly in the after-hours window where a human isn’t standing by to catch what the agent doesn’t flag.


The idea to test CaseGen came from a colleague. If you have a voice AI platform I should test, either yours or one you’re using, let me know.

Leave a Reply 0