
Neil Zeghidour
Co-founder and CEO of Gradium, a Paris-founded AI company building foundational audio models and low-latency infrastructure for real-time voice agents. Previously worked on generative-audio research at Meta and Google DeepMind and helped create Kyutai.
Full-Duplex Voice Agents Must Separate Conversation From Reasoning
Neil Zeghidour, co-founder and CEO of Gradium, argues that most real-time voice agents remain “walkie-talkies”: however quickly they respond, they can only listen or speak, not handle the overlapping signals that make human conversation work. Full-duplex speech models can represent simultaneous talk and backchanneling, he says, but they still sacrifice reasoning and tool-use capability relative to text-based agents. His proposed answer is a split architecture in which a lightweight speech model manages the live conversation while a separate text model handles harder reasoning and actions.
Voice AI Still Confuses Natural Speech With Real Conversation
Neil Zeghidour, CEO of Gradium AI and one of the researchers behind the full-duplex voice model Moshi, argues that voice AI’s long-promised “Her” moment is still being confused with better synthetic speech. His case is that cascaded voice agents are useful but structurally too slow and lossy to feel conversational, while speech-to-speech models improve flow but remain limited unless they can listen and speak simultaneously, use tools reliably, understand paralinguistic cues, and run cheaply enough to scale.