Subtitles and Trim
Auto-subtitles, speaker-aware lipsync, and video trim on upload
Two workflow upgrades that remove friction from video generation and editing: burn styled subtitles into any video automatically, and trim long clips before a generation runs.
Auto-Subtitles
Upload any video to the video editing tools and Kolbo can transcribe it and burn styled subtitles directly into the file, automatically, with no manual timing or captioning.
The subtitles are synced to the speech in the video. Style options control the appearance: font weight, size, placement, and color. The output is a single video file with the subtitles baked in and ready to post.
Useful for social content, educational clips, and any video where you want captions without the usual SRT-syncing workflow.
Speaker-Aware Lipsync
When you run a lipsync generation with multiple people on screen, you can pick which person gets the lip-sync treatment.
A visual selector appears over the video frame. Click the face you want to target, and the generation runs on that person specifically, leaving the other faces unchanged. This matters for interview-style footage, multi-character scenes, and any clip where you know exactly which speaker should be animated.
Video Trim on Upload
Most video generation models have a maximum input length. When you upload a clip that is longer than the selected model accepts, Kolbo surfaces a trim control directly in the upload flow. Set the in and out points for the range you want, and only that segment is sent to the model. No external editing step needed.
The same trim control also works proactively: even when your video is within limits, you can trim it to the specific range you want before the generation runs.
How to Use
- Auto-Subtitles: open the video editing tools, upload your video, and generate subtitles to burn them in.
- Speaker-aware lipsync: start a lipsync generation on multi-face footage and click the target face in the selector.
- Trim on upload: when a clip is too long, or whenever you want a specific range, set the in and out points before generating.
Tips
- Trimming to just the range you need keeps generations within model limits and avoids wasted processing.
- For speaker-aware lipsync, pick the face deliberately rather than relying on the model to guess the most prominent person.