IBM Research has identified a critical reliability problem plaguing AI agents: they fail to consistently reproduce successful task completions, even when tasked with identical requests. While ReAct agents powered by GPT-4.1 achieve a 77.4% success rate on average, they succeed on all five repeated runs for only 53.0% of tasks—a 24.4-percentage-point consistency gap that standard benchmarks obscure. This inconsistency stems not from capability limitations but from the shape of probability distributions underlying LLM decision-making. When an agent's decisions emerge from "flat" distributions—where multiple options carry nearly equal probability—minor platform-side fluctuations such as GPU floating-point variations or request batching effects can reorder which token gets selected, causing the agent to take a different path and fail on the next identical request.
To address this, IBM Research introduced Consistency Guidelines, a new diagnostic tool within the ALTK-Evolve system that identifies "flip-prone" decision points within recorded agent trajectories. The Consistency Analyzer resamples an agent's own recorded decisions to pinpoint vulnerable steps, then injects targeted guidelines to harden those decisions at inference time. Early results show the approach halves the consistency gap—dropping it from 24.4 to 12.0 percentage points—without sacrificing average accuracy. For mission-critical applications like financial reconciliation or contract review, this reliability gain is essential.
Key Points
Standard benchmarks report Mean@k (average success) while hiding Pass^k (success on all k repeated runs), obscuring a 24.4-point consistency gap in state-of-the-art ReAct agents
Inconsistency stems from 'flat' probability distributions in token selection, where small platform variations can flip downstream decisions across multi-step trajectories
IBM's Consistency Analyzer uses single-trajectory resampling to identify vulnerable decision points without re-running full tasks
Consistency Guidelines reduce the gap by 50% without costing average-case accuracy, addressing a reliability axis orthogonal to raw capability