🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Research Staff will develop foundational voice AI technologies, including low-bitrate neural audio codecs, steerable speech generation, disentangled audio representations, latent recombination, synthetic audio data generation, and multimodal speech-to-speech models. The role also involves designing scalable architectures, training methods, and inference algorithms optimized for hardware, billion-hour datasets, and real-time deployment. Candidates need strong mathematical foundations, foundation-model expertise, large-scale data pipeline experience, rigorous experimentation skills, deployment optimization knowledge, and publications or open-source contributions in speech or language AI.