Previous research

Speech systems that stay coherent, responsive, and efficient.

This body of work explores generative models for speech—from the representations that make them efficient to the alignment and conversational structure that make them reliable.

01

Long-form generation

Maintaining intelligibility, prosody, speaker identity, and natural boundaries across extended speech.

02

Speech & language

Connecting pretrained language models with audio representations for recognition, translation, and generation.

03

Robustness & control

Reducing speech hallucinations while improving control over speaker, accent, and delivery.

04

Efficient adaptation

Extending capable speech systems to new speakers and tasks without retraining every parameter.

Selected work in speech AI

Five projects, in context.

Each project addresses a different gap between impressive speech generation and systems that remain useful under real conditions.

01
Interspeech 2026Long-form TTSLead author

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form Data

Modern TTS models can sound natural for a sentence yet drift in prosody, speaker identity, or alignment across paragraphs. MagpieTTS-LF treats long-form generation as an inference problem rather than requiring a new model trained on long recordings.

The approach

Soft attention priors guide monotonic alignment while retaining past and future context. A stateful generation algorithm carries information across sentence chunks, and history-aware text encoding supports discourse-level prosodic planning.

Why it matters

The method extends an utterance-level model to coherent long-form synthesis without retraining on long-form data, improving continuity at both the speaker and sentence-boundary level.

02
Interspeech 2024Robust TTSAlignment

Improving Robustness of LLM-Based Speech Synthesis by Learning Monotonic Alignment

LLM-based speech models can repeat words, skip content, or lose alignment—especially when text contains repeated tokens. This work studies the cross-attention patterns behind those failures and makes the learned text-to-speech alignment more explicit.

The approach

Connectionist Temporal Classification loss and attention priors encourage monotonic cross-attention between text and speech tokens. The guidance improves alignment without adding learnable parameters to the model.

Why it matters

More reliable alignment directly reduces common speech hallucinations, improving transcript fidelity while preserving the flexibility of a large generative model.

03
ICML Audio Workshop 2025Preference alignmentControllable TTS

Koel-TTS: Enhancing LLM-Based Speech Generation with Preference Alignment and Classifier-Free Guidance

Autoregressive speech models offer expressive output but can produce unwanted vocalizations, weak transcript adherence, or a voice that drifts from the reference speaker. Koel-TTS brings alignment techniques from language modeling into speech generation.

The approach

Preference optimization uses feedback from speech-recognition and speaker-verification models. Classifier-free guidance further steers generation toward the requested transcript and reference voice.

Why it matters

The resulting system balances naturalness with control, improving intelligibility and speaker similarity while using a more focused training setup.

04
Interspeech 2023AdaptersSpeaker adaptation

Adapter-Based Extension of Multi-Speaker Text-to-Speech Models for New Speakers

Adding a new voice to a multi-speaker TTS system commonly requires expensive fine-tuning and can degrade voices the model already knows. This project isolates speaker adaptation from the core model.

The approach

Small adapter modules are inserted into a pretrained TTS network. The original model remains frozen while only the new adapters learn from a target speaker’s data.

Why it matters

Parameter-efficient adaptation preserves the shared base model, reduces training overhead, and makes it easier to maintain many customized voices.

05
ICASSP 2024Speech & languageIn-context learning

SALM: Speech-Augmented Language Model with In-Context Learning for Speech Recognition and Translation

Speech recognition and speech translation are often handled by separate task-specific systems. SALM explores a unified model that brings audio into a pretrained text language model while retaining its instruction-following and in-context capabilities.

The approach

An audio encoder and modality adapter connect speech input to a frozen text LLM. Lightweight LoRA layers accommodate speech and task instructions within a shared framework for recognition and translation.

Why it matters

A unified speech-language model can approach specialized baselines while supporting zero-shot in-context behavior, creating a more flexible foundation for speech tasks.

Patents & intellectual property

Applied ideas

2025

Speech-to-text processing assisted with language models

Conversational AI systems and applications · US Patent App. 18/408,400

2024

Synthetic speech generation

Conversational AI systems and applications · Publication US20240161728A1

2024

Customizing text-to-speech language models using adapters

Conversational AI systems and applications · Publication US20240127788A1

2018

Efficient task planning using past performance

US Patent US20180150786A1

Full publication record

Find citations and additional work on Google Scholar.

Open Scholar