Skip to navigation

Analyze images

This guide shows how to analyze one or more images with a prompt. Upload your images as assets and reference their identifiers, or pass them inline by URL or base64-encoded data. You can analyze up to twenty images in a single request.

On the Free plan, each image you analyze counts as 5 minutes toward a shared limit that also covers indexing. For example, 120 images use the entire 10-hour allowance.

For details on how your usage is measured and billed, see the Pricing page.

Key concepts

This section explains the key concepts and terminology used in this guide:

  • Asset: Your uploaded content. Once created, you can reference the same asset across multiple operations without uploading the file again.
  • Analysis task: An asynchronous operation for processing your content and generating text. Contains a status and the resulting text when complete.

Prerequisites

  • To use the platform, you need an API key:

    1

    If you don’t have an account, sign up for a free account.

    2

    Go to the API Keys page.

    3

    If you need to create a new key, select the Create API Key button. Enter a name and set the expiration period. The default is 12 months.

    4

    Select the Copy icon next to your key to copy it to your clipboard.

  • Depending on the programming language you are using, install the TwelveLabs SDK by entering one of the following commands:

    pip install --upgrade twelvelabs
  • Your images must meet the following requirements:

    • Number of images: one to twenty per request
    • Formats: JPEG, PNG, WebP, GIF, and BMP
    • File size: ≤ 20 MB per image
    • Pixel count: ≤ 16,777,216 pixels per image (width × height)

Complete example

Copy and paste the code below, replacing the placeholders surrounded by <> with your values. The example analyzes two images - the first uploaded as an asset, the second passed inline by URL - and references both in the prompt.

import time
from twelvelabs import TwelveLabs
from twelvelabs.types import AnalyzeImageInput
# 1. Initialize the client
client = TwelveLabs(api_key="<YOUR_API_KEY>")
# 2. Upload an image
asset = client.assets.create(
method="url",
url="<YOUR_FIRST_IMAGE_URL>" # Use direct links to raw media files. Video hosting platforms and cloud storage sharing links are not supported
# Or use method="direct" and file=open("<PATH_TO_IMAGE_FILE>", "rb") to upload a local file
)
print(f"Created asset: id={asset.id}")
# 3. Check the status of the asset
print("Waiting for asset to be ready...")
while True:
asset = client.assets.retrieve(asset.id)
if asset.status == "ready":
print("Asset is ready")
break
if asset.status == "failed":
raise RuntimeError(f"Asset processing failed: id={asset.id}")
time.sleep(5)
# 4. Analyze your images
images = [
AnalyzeImageInput(name="product_a", asset_id=asset.id),
AnalyzeImageInput(name="product_b", url="<YOUR_SECOND_IMAGE_URL>"), # Or upload it as an asset like the first image
]
task = client.analyze_async.tasks.create(
model_name="pegasus1.6",
image=images,
prompt="Compare <@product_a> with <@product_b> and describe what changed.",
# temperature=0.2,
# max_tokens=1024,
# You can also use `response_format` to request structured JSON responses
)
print(f"Task ID: {task.task_id}")
# 5. Monitor the status
while True:
task = client.analyze_async.tasks.retrieve(task.task_id)
if task.status == "ready":
print("Task completed")
break
elif task.status == "failed":
print("Task failed")
break
else:
print("Task still processing...")
time.sleep(5)
# 6. Process the results
print(f"{task.result.data}")

Code explanation

1

Import the SDK and initialize the client

Create a client instance to interact with the TwelveLabs Video Understanding Platform.
Function call: You call the constructor of the TwelveLabs class.
Parameters:

  • api_key: The API key to authenticate your requests to the platform.

Return value: An object of type TwelveLabs configured for making API calls.

2

Upload an image

Upload the first image to create an asset. Image analysis requires the asset to be in the ready state before you use it.
Function call: You call the assets.create method.
Parameters:

  • method: How to provide the image. Use url with a publicly accessible image URL, or direct with a local file.
3

Check the status of the asset

Asset processing is asynchronous. Poll the status of the asset until it is ready before you use it.

4

Analyze your images

Create an analysis task to start processing your images. This operation is asynchronous.
Function call: You call the analyze_async.tasks.create method.
Parameters:

  • image: A list of AnalyzeImageInput objects, one per image. You can analyze up to twenty images per request. For each image, set:

    • name: The name you use to reference the image in the prompt. Can contain only letters, digits, and underscores.
    • Exactly one source: asset_id, url, or base_64_string.

    This example uses the asset identifier from the previous step for the first image and passes the second image inline by URL.

  • prompt: The instructions for the analysis. To refer to a specific image in your prompt, use its name as the <@name> placeholder. Reference every image or none of them. If the prompt references none of them, the platform analyzes every image you provided. If it references only some of them, the platform returns a 400 error.

  • (Optional) temperature: Controls the randomness of the text output. A higher value generates more creative text, while a lower value produces more deterministic output.

  • (Optional) max_tokens: The maximum response length, in tokens.

  • (Optional) response_format: Use this parameter to request structured JSON responses. For instructions, examples, and best practices, see the Structured responses page.

Return value: An object of type CreateAnalyzeTaskResponse containing a field named task_id, which represents the unique identifier of your analysis task. You can use this identifier to track the status of your task.

5

Monitor the status

The platform requires some time to process images. Poll the status of the analysis task until processing completes. This example uses a loop to check the status every 5 seconds.
Function call: You repeatedly call the analyze_async.tasks.retrieve method until the task completes.

Parameters:

  • task_id: The unique identifier of your analysis task.

Return value: An object of type AnalyzeTaskResponse containing, among other information, the following fields:

  • status: The current status of the task. The possible values are:
    • queued: The task is waiting to be processed.
    • pending: The task is queued and waiting to start.
    • processing: The platform is analyzing your images.
    • ready: Processing is complete. Results are available in the result field.
    • failed: The task failed.
  • result: When the status is ready, this field contains the generated text and usage information.
6

Process the results

This example prints the generated text to the standard output.

Synchronous analysis

For immediate results, use the synchronous method instead of creating an analysis task. It supports the same image requirements and returns (or streams) the generated text directly in the response.

Response methods

Streaming responses deliver text fragments in real time as they are generated, enabling immediate processing and feedback. This method is the default behavior of the platform and is ideal for applications requiring incremental updates.

  • Response format: A stream of JSON objects in NDJSON format, with three event types:
    • stream_start: Marks the beginning of the stream.
    • text_generation: Delivers a fragment of the generated text.
    • stream_end: Signals the end of the stream.
  • Response handling:
    • Iterate over the stream to process text fragments as they arrive.
  • Advantages:
    • Real-time processing of partial results.
    • Reduced perceived latency.
  • Use case: Live captioning, real-time analysis, or applications needing instant updates.

Non-streaming responses deliver the complete generated text in a single response, simplifying processing when the full result is needed.

  • Response format: A single string containing the full generated text.
  • Response handling:
    • Access the complete text directly from the response.
  • Advantages:
    • Simplicity in handling the full result.
    • Immediate access to the entire text.
  • Use case: Generating reports, descriptions, or any scenario where the whole text is required at once.

Copy and paste the code below, replacing the placeholders surrounded by <> with your values.

from twelvelabs import TwelveLabs
from twelvelabs.types import AnalyzeImageInput
# 1. Initialize the client
client = TwelveLabs(api_key="<YOUR_API_KEY>")
# 2. Analyze your images
images = [
AnalyzeImageInput(name="product_a", url="<YOUR_FIRST_IMAGE_URL>"),
AnalyzeImageInput(name="product_b", url="<YOUR_SECOND_IMAGE_URL>"),
]
text_stream = client.analyze_stream(
model_name="pegasus1.6",
image=images,
prompt="Compare <@product_a> with <@product_b> and describe what changed.",
# temperature=0.2,
# max_tokens=1024,
# You can also use `response_format` to request structured JSON responses
)
# 3. Process the results
for text in text_stream:
if text.event_type == "text_generation":
print(text.text)

The image parameter and all optional parameters (model_name, temperature, max_tokens, response_format, etc.) function the same as in the asynchronous approach above.