Summarizing YouTube videos is a common transcript API use case: get the spoken text, chunk if needed, send it to an LLM with a prompt.
This walkthrough uses FreeTranscriptAPI and any OpenAI-compatible model. For endpoint details, see our developer guide.
Step 1: Extract the transcript
Pass a YouTube URL. The API returns JSON with text segments and timestamps. Check the API reference for query parameters.
Example request
curl "https://api.freetranscriptapi.com/v1/transcript?video_url=https://www.youtube.com/watch?v=dQw4w9WgXcQ"Step 2: Prepare text for the LLM
Join segments into one string. For long videos, chunk by time (say, 10-minute blocks), summarize each chunk, then summarize those summaries.
Python example
import requests
response = requests.get(
"https://api.freetranscriptapi.com/v1/transcript",
params={"video_url": "https://www.youtube.com/watch?v=VIDEO_ID"},
)
data = response.json()
full_text = " ".join(segment["text"] for segment in data["transcript"])
summary_prompt = f"Summarize this video transcript in 3 bullet points:\n\n{full_text}"Step 3: Send to your LLM
Pass the text to GPT-4, Claude, Gemini, or whatever you run. The model reads captions, not audio.
Why timestamps matter
Each segment has start and duration in seconds. That enables:
- Timestamped summaries ("At 5:30, the speaker discusses...")
- Chapter generation from topic shifts
- Linking summary bullets back to video moments
- Re-summarizing only part of a long video
Handling long videos
A 2-hour podcast can hit 30,000+ words. Most models have context limits, so map-reduce:
- Split the transcript by time or word count
- Summarize each chunk
- Summarize the chunk summaries
With 1,000 signup credits, you can run a lot of videos through this before buying more for the transcript step. Compare free tiers in our free API roundup.
Automation with n8n
Chain an HTTP Request node (FreeTranscriptAPI) with an OpenAI node. Trigger on new URLs from a sheet, RSS feed, or webhook. No custom code required.