AssemblyAI runs speech-to-text on audio and video files. FreeTranscriptAPI pulls captions YouTube already has. Pick the wrong one and you pay for inference you did not need. Our API guide covers when caption extraction is enough.
Caption extraction vs speech-to-text
| Factor | FreeTranscriptAPI | AssemblyAI |
|---|---|---|
| Input | YouTube URL or video ID | Audio/video file upload or URL |
| Method | Extracts existing captions | AI speech recognition |
| Speed | Under 2 seconds | Real-time to 2x audio length |
| Cost per video | Free tier available | ~$0.15-0.37/hour of audio |
| Accuracy source | YouTube creator captions | AI model inference |
| Timestamps | Per-line JSON | Word-level timestamps |
When FreeTranscriptAPI fits
- The video already has captions (most do)
- You need cheap extraction at volume
- You want the creator's captions, not a model guess
- Your input is a YouTube URL, not an audio file
- You want a free tier without a credit card (see our free tier comparison)
When AssemblyAI fits
- No captions exist and you need AI transcription
- You need word-level timestamps for editing
- You are transcribing podcasts, meetings, or calls
- Speaker diarization matters
The hybrid approach
Many pipelines try caption extraction first (fast, often free), then call speech-to-text only when captions are missing. FreeTranscriptAPI covers the first step. See the API docs for request examples.
Example request
curl "https://api.freetranscriptapi.com/v1/transcript?video_url=https://www.youtube.com/watch?v=dQw4w9WgXcQ"