✨ From vibe coding to vibe deployment. UBOS MCP turns ideas into infra with one message.

Learn more
Andrii Bidochko
  • Updated: March 26, 2026
  • 2 min read

Tencent AI Launches Covo‑Audio 7B: Real‑Time Speech‑Language Model Redefining Conversational AI

Tencent AI Unveils Covo‑Audio 7B – A Game‑Changing Real‑Time Speech‑Language Model

In a bold move that could reshape the future of conversational AI, Tencent AI has open‑sourced Covo‑Audio 7B, a 7‑billion‑parameter speech‑language model designed for low‑latency, high‑fidelity audio interactions.

The model introduces a unified end‑to‑end architecture that interleaves three modalities—audio waveforms, token embeddings, and textual tokens—in a hierarchical tri‑modal fashion. This design enables intelligence‑speaker decoupling, allowing the system to reason independently of the speaker’s voice and deliver full‑duplex, real‑time chat capabilities.

Key Highlights:

  • Unified Architecture: Seamlessly merges speech recognition, language understanding, and speech synthesis in a single pipeline.
  • Hierarchical Tri‑Modal Interleaving: Aligns audio, token, and text streams for more coherent and context‑aware responses.
  • Full‑Duplex Interaction: Supports simultaneous speaking and listening, mimicking natural human conversation.
  • Efficient Training: Trained on a mixture of public and proprietary datasets, achieving strong performance on benchmarks such as MMAU, MMSU, URO‑Bench, and VStyle while keeping inference costs low.

The open‑source release includes the model weights, inference pipeline, and a set of evaluation scripts, inviting developers and researchers to experiment, fine‑tune, and integrate Covo‑Audio into a variety of applications—from virtual assistants and customer support bots to real‑time translation services.

For those looking to dive deeper into speech‑language technologies, explore our Speech Model Resource Center and stay updated with the latest AI breakthroughs on our AI News Hub.

With Covo‑Audio 7B, Tencent AI not only pushes the boundaries of multimodal alignment but also sets a new standard for real‑time, interactive AI experiences. The model’s open‑source nature promises rapid community adoption and innovation, potentially accelerating the next wave of conversational AI solutions.


Andrii Bidochko

CTO UBOS

Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.

Sign up for our newsletter

Stay up to date with the roadmap progress, announcements and exclusive discounts feel free to sign up with your email.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.