Long-form generation
Maintaining intelligibility, prosody, speaker identity, and natural boundaries across extended speech.
Previous research
This body of work explores generative models for speech—from the representations that make them efficient to the alignment and conversational structure that make them reliable.
Maintaining intelligibility, prosody, speaker identity, and natural boundaries across extended speech.
Connecting pretrained language models with audio representations for recognition, translation, and generation.
Reducing speech hallucinations while improving control over speaker, accent, and delivery.
Extending capable speech systems to new speakers and tasks without retraining every parameter.
Selected work in speech AI
Each project addresses a different gap between impressive speech generation and systems that remain useful under real conditions.
Modern TTS models can sound natural for a sentence yet drift in prosody, speaker identity, or alignment across paragraphs. MagpieTTS-LF treats long-form generation as an inference problem rather than requiring a new model trained on long recordings.
Soft attention priors guide monotonic alignment while retaining past and future context. A stateful generation algorithm carries information across sentence chunks, and history-aware text encoding supports discourse-level prosodic planning.
The method extends an utterance-level model to coherent long-form synthesis without retraining on long-form data, improving continuity at both the speaker and sentence-boundary level.
LLM-based speech models can repeat words, skip content, or lose alignment—especially when text contains repeated tokens. This work studies the cross-attention patterns behind those failures and makes the learned text-to-speech alignment more explicit.
Connectionist Temporal Classification loss and attention priors encourage monotonic cross-attention between text and speech tokens. The guidance improves alignment without adding learnable parameters to the model.
More reliable alignment directly reduces common speech hallucinations, improving transcript fidelity while preserving the flexibility of a large generative model.
Autoregressive speech models offer expressive output but can produce unwanted vocalizations, weak transcript adherence, or a voice that drifts from the reference speaker. Koel-TTS brings alignment techniques from language modeling into speech generation.
Preference optimization uses feedback from speech-recognition and speaker-verification models. Classifier-free guidance further steers generation toward the requested transcript and reference voice.
The resulting system balances naturalness with control, improving intelligibility and speaker similarity while using a more focused training setup.
Adding a new voice to a multi-speaker TTS system commonly requires expensive fine-tuning and can degrade voices the model already knows. This project isolates speaker adaptation from the core model.
Small adapter modules are inserted into a pretrained TTS network. The original model remains frozen while only the new adapters learn from a target speaker’s data.
Parameter-efficient adaptation preserves the shared base model, reduces training overhead, and makes it easier to maintain many customized voices.
Speech recognition and speech translation are often handled by separate task-specific systems. SALM explores a unified model that brings audio into a pretrained text language model while retaining its instruction-following and in-context capabilities.
An audio encoder and modality adapter connect speech input to a frozen text LLM. Lightweight LoRA layers accommodate speech and task instructions within a shared framework for recognition and translation.
A unified speech-language model can approach specialized baselines while supporting zero-shot in-context behavior, creating a more flexible foundation for speech tasks.
Patents & intellectual property
Conversational AI systems and applications · US Patent App. 18/408,400
Conversational AI systems and applications · Publication US20240161728A1
Conversational AI systems and applications · Publication US20240127788A1
US Patent US20180150786A1
Full publication record