Rather than building speech infrastructure from scratch, JotPsych's technical team made an early strategic decision to integrate best-in-class transcription technology, allowing them to launch quickly while maintaining focus on their clinical workflow innovations.
Build vs. buy decision framework
The team evaluated the opportunity cost of building speech-to-text capabilities versus integrating an API that would let them focus entirely on behavioral health-specific features: clinical note generation, terminology handling, and workflow optimization.
Technical requirements for medical context
JotPsych needed several key capabilities: exceptionally high transcription accuracy for medical terminology, robust speaker diarization to handle multi-participant sessions, real-time streaming for immediate clinician feedback, and reliable performance across varied clinical environments.
Implementation timeline priorities
As a startup entering a competitive market, speed-to-market was critical. The team needed a solution they could implement immediately and build upon incrementally as they refined their product offering.
AssemblyAI has been a fundamental block to our business, enabling us to focus product efforts on behavioral health-specific workflows.
Jackson Bierfeldt
Co-founder & CTO, JotPsych
The company selected AssemblyAI's speech-to-text APIs, including Universal Speech-to-Text for pre-recorded batch transcription, Speech Understanding models like Speaker Diarization and PII redaction, and Universal-Streaming Speech-to-Text for real-time applications.