Drop Xiulan into Claude and get a speech engineer who reports word error rate by accent and condition, because a single headline WER hides exactly the users you are failing. Xiulan owns the audio modality end to end: automatic speech recognition with Whisper and its successors, streaming versus batch decoding, WER measurement and what it conceals, domain adaptation, custom vocabulary and contextual biasing; speaker diarization and speaker identification; voice activity detection; text to speech and voice cloning under consent constraints; audio preprocessing including resampling, noise reduction, echo cancellation and channel separation; keyword spotting and wake words; audio classification and event detection; real-time streaming pipelines with latency budgets; and the fairness problem of recognition quality varying sharply by accent and dialect. What you get →ASR tuning: domain adaptation, custom vocabulary, biasing →WER measured by accent, noise condition and speaker, not headline →Diarization, speaker ID, VAD and channel separation →Streaming pipelines, latency budgets, TTS and consent limits 📄 xiulan-speech-audio-ml-engineer.skill Under 2 min install Works with Claude, ChatGPT & any AI chat How to install Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Xiulan builds the answer. Includes a full worked example so you see exactly what you get.