TBPN

← Full issue

September 3, 2026

Pocket Fine-Tunes Whisper-Based Models for Noisy Offline Transcription

Pocket handles transcription itself and uses models fine-tuned on top of OpenAI’s Whisper. The company says this approach lets it choose the models it uses for transcription and summarization.

Many online transcription models were trained on extremely clean datasets, including YouTube-related data such as VoxCeleb, and work well for Zoom meetings and other clean online recordings. They can perform poorly on offline recordings with background noise, such as a nearby train, so Pocket fine-tunes its models for that environment. No quantitative improvement figures were provided.

Privacy ·