CfP | RTCA @ NeurIPS 2026 [R]

· Source: Machine Learning · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Expert, quick

Summary

The inaugural Real-Time Conversational Agents (RTCA) Workshop at NeurIPS 2026, scheduled for 11 or 12 December 2026 in Sydney, Australia, has issued a Call for Papers and Demos. This workshop focuses on advancing real-time multimodal conversational agents, addressing challenges in streaming speech, video, and language generation, achieving naturalness in interaction, and evaluating live systems. The field aims to move beyond offline generation to systems that operate in real time, handling latency, turn-taking, backchannels, and cross-modal alignment. Contributions are invited on topics such as streaming/low-latency speech synthesis, real-time talking-head generation, incremental language models, multimodal alignment, and evaluation of naturalness. Submissions include full papers (up to 8 pages), short papers (up to 4 pages), and demo papers (up to 2 pages), all requiring the NeurIPS 2026 style file and double-blind review via OpenReview. The submission deadline is 29 August 2026, with author notifications by 29 September 2026. The workshop is non-archival.

Key takeaway

For AI Scientists and NLP Engineers developing conversational AI, this Call for Papers highlights critical research gaps in real-time multimodal interaction. You should consider submitting original work on streaming models, turn-taking, or live system evaluation to the RTCA Workshop at NeurIPS 2026. This is an opportunity to shape emerging benchmarks and methodologies for natural, low-latency agents, contributing to a non-archival venue that allows subsequent publication elsewhere.

Key insights

Real-time multimodal conversational agents require new approaches for natural interaction and live system evaluation.

Principles

Method

The workshop seeks contributions on real-time generation under latency, naturalness in interaction, and evaluation of live systems, covering streaming models, multimodal alignment, and interaction handling.

In practice

Topics

Best for: Research Scientist, AI Scientist, NLP Engineer, Computer Vision Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Machine Learning.