IBM Research has identified a critical reliability problem plaguing AI agents: they fail to consistently reproduce successful task completions, even when tasked with identical requests. While ReAct agents powered by GPT-4.1 achieve a 77.4% success rate on average, they succeed on all five repeated runs for only 53.0% of tasks—a 24.4-percentage-point consistency gap that standard benchmarks obscure. This inconsistency stems not from capability limitations but from the shape of probability distributions underlying LLM decision-making. When an agent's decisions emerge from "flat" distributions—where multiple options carry nearly equal probability—minor platform-side fluctuations such as GPU floating-point variations or request batching effects can reorder which token gets selected, causing the agent to take a different path and fail on the next identical request. To address this, IBM Research introduced Consistency Guidelines, a new diagnostic tool within the ALTK-Evolve system that identifies "flip-prone" decision points within recorded agent trajectories. The Consistency Analyzer resamples an agent's own recorded decisions to pinpoint vulnerable steps, then injects targeted guidelines to harden those decisions at inference time. Early results show the approach halves the consistency gap—dropping it from 24.4 to 12.0 percentage points—without sacrificing average accuracy. For mission-critical applications like financial reconciliation or contract review, this reliability gain is essential.