Analyze videos and images
The platform uses a multimodal approach to analyze videos and images and generate text, processing visuals, sounds, spoken words, and on-screen text to provide a comprehensive understanding. This method captures nuances that unimodal interpretations might miss, enabling accurate, context-rich text generation based on your content.
Key features:
- Multimodal analysis: Processes visuals, sounds, spoken words, and text for a holistic understanding of your content.
- Egocentric video understanding: Analyzes first-person footage from sources such as wearable cameras and teleoperated systems.
- Image analysis: Analyzes images with the same prompt-based workflow as videos.
- Entity recognition: Identifies and names people, objects, and characters in a scene.
- Customizable prompts: Allows tailored outputs through instructive, descriptive, or question-based prompts.
- Flexible text generation: Supports various tasks, including summarization, chaptering, and open-ended text generation.
- Segment videos: Extract structured, timestamped segments from your videos by defining custom segment types and fields, including lists of timestamped events inside each segment.
- Segmentation without a token limit: Segments long videos and returns the segments it produced, even when the output is incomplete.
Use cases:
- Robotics and physical AI: Label actions, check quality, and predict teleoperated trajectories and next actions in egocentric footage.
- Content structuring: Organize and structure content for e-learning platforms to improve usability.
- Product analysis: Compare product images or extract product details and visible features.
- SEO optimization: Optimize content to rank higher in search engine results.
- Highlight creation: Create short, engaging video clips for media and broadcasting.
- Incident reporting: Record and report incidents for security and law enforcement purposes.
- Generative video QA: Check AI-generated clips against their prompt or flag visual defects before you use them.
For details on how your usage is measured and billed, see the Pricing page.
Workflow
The platform provides two methods to analyze content. Choose the method that fits your use case:
Customize text generation
You can configure the temperature to control output randomness, set the maximum response length, and request structured JSON responses for programmatic processing. To extract timestamped segments with custom fields, see the Segment videos page.