CfP | RTCA @ NeurIPS 2026 [R]
Summary
The inaugural Real-Time Conversational Agents (RTCA) Workshop at NeurIPS 2026, scheduled for 11 or 12 December 2026 in Sydney, Australia, has issued a Call for Papers and Demos. This workshop focuses on advancing real-time multimodal conversational agents, addressing challenges in streaming speech, video, and language generation, achieving naturalness in interaction, and evaluating live systems. The field aims to move beyond offline generation to systems that operate in real time, handling latency, turn-taking, backchannels, and cross-modal alignment. Contributions are invited on topics such as streaming/low-latency speech synthesis, real-time talking-head generation, incremental language models, multimodal alignment, and evaluation of naturalness. Submissions include full papers (up to 8 pages), short papers (up to 4 pages), and demo papers (up to 2 pages), all requiring the NeurIPS 2026 style file and double-blind review via OpenReview. The submission deadline is 29 August 2026, with author notifications by 29 September 2026. The workshop is non-archival.
Key takeaway
For AI Scientists and NLP Engineers developing conversational AI, this Call for Papers highlights critical research gaps in real-time multimodal interaction. You should consider submitting original work on streaming models, turn-taking, or live system evaluation to the RTCA Workshop at NeurIPS 2026. This is an opportunity to shape emerging benchmarks and methodologies for natural, low-latency agents, contributing to a non-archival venue that allows subsequent publication elsewhere.
Key insights
Real-time multimodal conversational agents require new approaches for natural interaction and live system evaluation.
Principles
- Real-time interaction is harder than offline generation.
- Latency and turn-taking are first-class problems.
- Interactional naturalness lacks shared benchmarks.
Method
The workshop seeks contributions on real-time generation under latency, naturalness in interaction, and evaluation of live systems, covering streaming models, multimodal alignment, and interaction handling.
In practice
- Explore full-duplex speech-language models.
- Develop real-time talking-head generation.
- Design interactive evaluation datasets.
Topics
- Real-Time Conversational Agents
- Multimodal Interaction
- NeurIPS 2026
- Streaming Language Models
- Speech Synthesis
- Talking-Head Generation
- Live System Evaluation
Best for: Research Scientist, AI Scientist, NLP Engineer, Computer Vision Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Machine Learning.