Teaching ASR to keep the words that matter Entity- and disfluency-aware speech recognition for accented conversational English. INTERSPEECH 2026. How I evaluate a voice agent What to measure in a speech-to-speech system when there is no reference transcript to compare against. Serving speech models when the metrics lie Why GPU utilisation is the wrong signal for multi-model inference pipelines, and what to instrument instead. Building an LLM judge you can actually trust Validating automated evaluation against human labels, and what to do when they disagree.