Two weeks ago your AI agent worked. Perfect, actually. Every test passed. You deployed it. Users loved it.
Yesterday it started generating broken SQL queries. Wrong answers. Plausible nonsense. You panic. Re-run your tests. Still 100% pass rate.
What happened?
Talk to anyone shipping agents in production and you'll hear this story. Different teams, different apps, same confusion. The tests say everything works. Production says otherwise. And the gap between those two realities costs weeks of debugging.
What Your Tests Missed
One team tested their SQL agent like this: "Show me Q3 revenue." Perfect query. "Who were our top customers?" Nailed it. "Compare revenue by region." Flawless. One hundred tests. One hundred passes.
But that's not how users talked to it.
Real usage looked like this:
"Show me Q3 revenue." "Now show me Q2." "What about last year?" "Wait, which table was that from again?"
Ten turns in, the agent drowns in context. It pulls table names from old error messages. Mixes up which quarter the user meant. Loses track of what "that" refers to. The conversation history, full of corrections and failed attempts, confuses it.
Your test suite tested questions. Production is conversations.
Why This Keeps Happening
A normal function is stable. Call calculate_tax(100000) and you get 25,000 every time. Test it once, trust it forever.
An agent isn't a function. It's a conversation system. Same question, different answer, depending on what happened ten minutes ago. The component you're testing isn't the prompt. It's the entire context.
And context grows. Error messages pile up. Clarifications stack. By turn fifteen, your agent interprets new questions through the lens of everything that came before. Good luck writing test cases for that.
Traditional software fails in ways you can isolate. Null pointer. Off-by-one error. Race condition. You fix them. You move on.
Agents fail in context. The same prompt works Monday. Fails Tuesday. The difference is the conversation that led up to it. You can't isolate the failure because the failure IS the accumulated state.
What Works Now
Test conversations, not questions. Stop running isolated test cases. String them together. "Show revenue for Q3" → "now Q2" → "what about last year?" If your test suite doesn't include multi-turn flows where later questions reference earlier ones, you're not testing reality.
Test with garbage in the context. Don't just test happy paths. Test what happens when the history includes three failed attempts, two corrections, and an unclear request. Because that's what your agent will see in production.
Version lock everything. Model version, not "latest." API version. Prompt templates. Retrieval logic. When behavior changes, you need to know what shifted. Debugging "something somewhere changed" is hell.
Log the full context, not just the final prompt. Production failures make no sense when you only see the last question. Log the entire conversation. Log what the agent retrieved. Log the timestamp and model version. Build the ability to replay exact production scenarios in your test environment.
Watch for drift like you'd watch for memory leaks. Your agent's behavior in week one won't match week eight. Users learn patterns. Conversations get longer. Error types shift. Track output distributions over time. Alert when patterns change.
The Truth
Your agent isn't broken. Your testing approach is.
You're testing like it's a function. Input A produces output B. But it's a stateful system where every interaction affects every future interaction. The traditional software testing playbook doesn't apply.
This doesn't mean agents can't work. It means the bar for "working" is higher than you thought. And "passed all tests" doesn't mean what it used to mean.
The teams who make this work aren't the ones with the best models. They're the ones who recognized this gap early and built testing infrastructure for conversation systems, not functions.
You can do this. But not with your current test suite.
For deeper reading: Anthropic's work on constitutional AI and multi-turn evaluation, Hamel Husain's writing on LLM testing frameworks, and the growing research on conversational benchmarks all point to this same gap between single-turn evals and production reality.
