Google researchers have introduced AgentHands, an LLM-powered prototype that equips AI conversational agents with synchronized, expressive hand gestures in Extended Reality (XR) environments. Published at CHI 2026, the system bridges the gap between abstract verbal instructions and intuitive physical demonstrations, allowing AI agents to gesture naturally while providing spatially grounded guidance. This addresses a key limitation of current AI assistants: while tools like Gemini 3.1 Flash Live can reference physical objects via 2D overlays, AgentHands enables truly embodied interactions in immersive platforms like Android XR.
The system uses a multi-dimensional taxonomy to define how virtual hands should interact with physical space. It categorizes gestures into handedness and gesture type (palm for caution, cylindrical grip for tool-holding), spatiality (mid-air, object-anchored, or user-relative positioning), temporal dynamics (animated motions like pouring or tracing), and visual effects like heat warnings. The workflow combines environment awareness through eye-gaze object registration, a gesture event library, LLM-based gesture reasoning, and synchronized text-to-speech execution that coordinates hand animations with spoken words at the word level.
Google demonstrated AgentHands in real-world scenarios including interactive tutoring—where an agent gestures to point out aerial roots while explaining orchid care—and technical walkthroughs for 3D printer operations. By combining the spatial reasoning of LLMs with XR's depth and motion capabilities, AgentHands transforms AI assistance from abstract text-based guidance into embodied, spatially aware demonstrations that users can intuitively follow and understand.
Key Points
AgentHands is an LLM-powered XR prototype that synchronizes AI agent hand gestures with speech for spatially grounded conversations
The system uses a taxonomy of gesture types (deictic, iconic, expressive) mapped to specific spatial positions and visual effects in XR
Core workflow includes environment awareness via eye-gaze, gesture event libraries, LLM-based reasoning, and word-level synchronized animation execution
Demonstrated applications show enhanced user engagement in interactive tutoring and technical walkthroughs by embodying abstract instructions as physical gestures