API ResourcesImages

Images

The Images API provides a comprehensive set of AI-powered image capabilities, including image generation, image editing, OCR text recognition, and image description. With these APIs, you can generate images from text, modify existing images, extract text from images, and even describe image content using natural language.


POST/beta/v1/images/generations

Image Generation

The Image Generation API is the core of text-to-image capabilities. Simply provide a detailed text description, and the model will create original images that match your description. This is ideal for building art generators, marketing poster designs, and product prototype visualizations.

Request Headers

  • Name
    Content-Type
    Type
    string
    Required
    Required
    Description
    Value must be application/json.
  • Name
    Authorization
    Type
    string
    Required
    Required
    Description
    Optional authentication method. The credential used for API authentication. For details, see Authentication.

Request Body

  • Name
    model
    Type
    string
    Required
    Required
    Description
    The model ID used for image generation. Please choose the appropriate model based on your needs. Currently supported models include gpt-image-1 and gemini-2.5-flash-image-preview.
  • Name
    prompt
    Type
    string
    Required
    Required
    Description
    The text description of the desired image. gpt-image-1 has a maximum length of 32000 characters, and dall-e-3 has a maximum of 4000 characters.
  • Name
    background
    Type
    string
    Optional
    Optional
    Description
    Allows setting background transparency for the generated image. This parameter only applies to gpt-image-1. Must be one of transparent, opaque, or auto (default). When using auto, the model will automatically determine the best background for the image.If transparent is selected, the output format must support transparency, so it should be set to png (default) or webp.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    moderation
    Type
    string
    Optional
    Optional
    Description
    Controls the content moderation level of images generated by the model. Must be set to low (less filtering restrictions) or auto (default).
  • Name
    n
    Type
    string
    Optional
    Optional
    Description
    The number of images to generate. Must be between 1 and 10. For dall-e-3 and gemini-2.5-flash-image-preview, only n=1 is supported.
  • Name
    output_compression
    Type
    integer
    Optional
    Optional
    Description
    The compression level of the generated image (0-100%). This parameter only applies to the gpt-image-1 model with output format webp or jpeg. Default is 100.
  • Name
    output_format
    Type
    string
    Optional
    Optional
    Description
    The format of the generated image returned. This parameter only applies to gpt-image-1. Must be one of png (default), jpeg, or webp.
  • Name
    partial_images
    Type
    string
    Optional
    Optional
    Description
    The number of partial images to generate. This parameter is used to return a streaming response with partial images. The value must be between 0 and 3. When set to 0, the response will be sent as a single image in one streaming event.Note: If the complete image is generated faster, the final image may be sent before all partial images are generated.
  • Name
    quality
    Type
    string
    Optional
    Optional
    Description
    The quality of the generated image.
    • auto(默认值)将自动为给定模型选择最佳质量。
    • gpt-image-1 supports high, medium, and low quality levels.
    • dall-e-3 supports hd and standard quality levels.
    Note: This parameter will be ignored for the gemini-2.5-flash-image-preview model.
  • Name
    response_format
    Type
    string
    Optional
    Optional
    Description
    The return format of images generated by DALL-E 3. Must be one of url or b64_json. URLs are valid for only 60 minutes after image generation. This parameter does not apply to gpt-image-1 and gemini-2.5-flash-image-preview, which always return base64 encoded images.
  • Name
    size
    Type
    string
    Optional
    Optional
    Description
    The size of the generated image. For gpt-image-1, must be one of 1024x1024, 1536x1024 (landscape), 1024x1536 (portrait), or auto (default); for dall-e-3, it is 1024x1024, 1792x1024, or 1024x1792.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    stream
    Type
    string
    Optional
    Optional
    Description
    Generate images in streaming mode. Default is false. This parameter does not apply to dall-e-3.
  • Name
    style
    Type
    string
    Optional
    Optional
    Description
    The style of the generated image. This parameter only applies to the dall-e-3 model. Must choose one of vivid or natural. Choosing "vivid" will make the model tend to generate hyper-realistic and dramatic images; choosing "natural" will make the model generate more natural, less hyper-realistic images.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    user
    Type
    string
    Optional
    Optional
    Description
    A unique identifier representing the end user, used for monitoring and detecting abuse.

Response Body

  • Name
    background
    Type
    string
    Required
    Required
    Description
    The background parameter used for image generation. Value is transparent or opaque.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    created
    Type
    integer
    Required
    Required
    Description
    Unix timestamp (in seconds) when the image was created.
  • Name
    data
    Type
    array
    Required
    Required
    Description
    List of generated images.
  • Name
    output_format
    Type
    string
    Required
    Required
    Description
    Output format for image generation. Values are png, webp, or jpeg.
  • Name
    quality
    Type
    string
    Required
    Required
    Description
    The quality of the generated image. Values are low, medium, or high.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    size
    Type
    string
    Required
    Required
    Description
    The size of the generated image. Values are 1024x1024, 1024x1536, or 1536x1024.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    usage
    Type
    object
    Required
    Required
    Description
    Token usage information for image generation. This parameter is not supported for the dall-e-3 model.

Request

POST
/beta/v1/images/generations
curl https://api.easytransnote.com/beta/v1/images/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $YOUR_API_KEY" \
  -d '{
    "model": "gpt-image-1",
    "prompt": "A cute baby sea otter",
    "n": 1,
    "size": "1024x1024"
  }'

Response

{
  "created": 1757908389,
  "data": [
    {
      "b64_json": "iVBORw0KGgoAAAANSUhEUgA"
    }
  ],
  "usage": {
    "input_tokens": 75,
    "input_tokens_details": {
      "image_tokens": 0,
      "text_tokens": 75
    },
    "output_tokens": 1024,
    "total_tokens": 1099
  }
}

POST/beta/v1/images/edits

Image Edit

The Image Edit API provides powerful image restoration and modification capabilities. Unlike generating from scratch, this API allows you to upload an original image and precisely modify specific areas through text instructions and optional masks. This API is widely used in photo restoration, product image refinement, creative composition, and other scenarios.

Request Headers

  • Name
    Content-Type
    Type
    string
    Required
    Required
    Description
    Value must be multipart/form-data.
  • Name
    Authorization
    Type
    string
    Required
    Required
    Description
    Optional authentication method. The credential used for API authentication. For details, see Authentication.

Request Body

  • Name
    model
    Type
    string
    Required
    Required
    Description
    The model ID used for image generation. Please choose the appropriate model based on your needs. Currently supported models include gpt-image-1 and gemini-2.5-flash-image-preview.
  • Name
    prompt
    Type
    string
    Required
    Required
    Description
    The text description of the desired image. gpt-image-1 has a maximum length of 32000 characters.
  • Name
    image
    Type
    file | array
    Required
    Required
    Description
    The image file to be edited. Must be a supported image format or an array of images.For gpt-image-1, each image should be in png, webp, or jpg format and less than 50MB. Up to 16 images can be provided.
  • Name
    background
    Type
    string
    Optional
    Optional
    Description
    Allows setting background transparency for the generated image. This parameter only applies to gpt-image-1. Must be one of transparent, opaque, or auto (default). When using auto, the model will automatically determine the best background for the image.If transparent is selected, the output format must support transparency, so it should be set to png (default) or webp.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    input_fidelity
    Type
    string
    Optional
    Optional
    Description
    Controls how much effort the model puts into matching the style and features (especially facial features) of the input image. This parameter only applies to the gpt-image-1 model. Supports high and low settings, default is low.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    mask
    Type
    file
    Optional
    Optional
    Description
    An additional image where completely transparent areas (e.g., where the Alpha channel value is zero) indicate the areas of the image that need to be edited. If multiple images are provided, the mask will be applied to the first image. Must be a valid PNG file, less than 4MB, and the same dimensions as the target image.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and additional mask images, and cannot guarantee the final image has this effect.
  • Name
    n
    Type
    string
    Optional
    Optional
    Description
    The number of images to generate. Must be between 1 and 10. For dall-e-3 and gemini-2.5-flash-image-preview, only n=1 is supported.
  • Name
    output_compression
    Type
    integer
    Optional
    Optional
    Description
    The compression level of the generated image (0-100%). This parameter only applies to the gpt-image-1 model with output format webp or jpeg. Default is 100.
  • Name
    output_format
    Type
    string
    Optional
    Optional
    Description
    The format of the generated image returned. This parameter only applies to gpt-image-1. Must be one of png (default), jpeg, or webp.
  • Name
    partial_images
    Type
    string
    Optional
    Optional
    Description
    The number of partial images to generate. This parameter is used to return a streaming response with partial images. The value must be between 0 and 3. When set to 0, the response will be sent as a single image in one streaming event.Note: If the complete image is generated faster, the final image may be sent before all partial images are generated.
  • Name
    quality
    Type
    string
    Optional
    Optional
    Description
    The quality of the generated image.
    • auto(默认值)将自动为给定模型选择最佳质量。
    • gpt-image-1 supports high, medium, and low quality levels.
    • dall-e-3 supports hd and standard quality levels.
    Note: This parameter will be ignored for the gemini-2.5-flash-image-preview model.
  • Name
    response_format
    Type
    string
    Optional
    Optional
    Description
    The return format of images generated by DALL-E 3. Must be one of url or b64_json. URLs are valid for only 60 minutes after image generation. This parameter does not apply to gpt-image-1 and gemini-2.5-flash-image-preview, which always return base64 encoded images.
  • Name
    size
    Type
    string
    Optional
    Optional
    Description
    The size of the generated image. For gpt-image-1, must be one of 1024x1024, 1536x1024 (landscape), 1024x1536 (portrait), or auto (default); for dall-e-3, it is 1024x1024, 1792x1024, or 1024x1792.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
  • Name
    stream
    Type
    string
    Optional
    Optional
    Description
    Generate images in streaming mode. Default is false.
  • Name
    user
    Type
    string
    Optional
    Optional
    Description
    A unique identifier representing the end user, used for monitoring and detecting abuse.

Response Body

  • Name
    background
    Type
    string
    Required
    Required
    Description
    The background parameter used for image generation. Value is transparent or opaque.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    created
    Type
    integer
    Required
    Required
    Description
    Unix timestamp (in seconds) when the image was created.
  • Name
    data
    Type
    array
    Required
    Required
    Description
    List of generated images.
  • Name
    output_format
    Type
    string
    Required
    Required
    Description
    Output format for image generation. Values are png, webp, or jpeg.
  • Name
    quality
    Type
    string
    Required
    Required
    Description
    The quality of the generated image. Values are low, medium, or high.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    size
    Type
    string
    Required
    Required
    Description
    The size of the generated image. Values are 1024x1024, 1024x1536, or 1536x1024.Note: For model compatibility, gemini-2.5-flash-image-preview also supports this parameter, and the corresponding value is always auto.
  • Name
    usage
    Type
    object
    Required
    Required
    Description
    Token usage information for image generation. This parameter is not supported for the dall-e-3 model.

Request

POST
/beta/v1/images/edits
curl -s -D >(grep -i x-request-id >&2) \
  -o >(jq -r '.data[0].b64_json' | base64 --decode > gift-basket.png) \
  -X POST "https://api.easytransnote.com/beta/v1/images/edits" \
  -H "Authorization: Bearer $YOUR_API_KEY" \
  -F "model=gpt-image-1" \
  -F "image[][email protected]" \
  -F "image[][email protected]" \
  -F "image[][email protected]" \
  -F "image[][email protected]" \
  -F 'prompt=Create a lovely gift basket with these four items in it'

Response

{
  "created": 1757908389,
  "data": [
    {
      "b64_json": "iVBORw0KGgoAAAANSUhEUgA"
    }
  ],
  "usage": {
    "input_tokens": 75,
    "input_tokens_details": {
      "image_tokens": 0,
      "text_tokens": 75
    },
    "output_tokens": 1024,
    "total_tokens": 1099
  }
}

POST/beta/v1/images/ocr

Image OCR

The Image OCR API is a powerful interface designed to accurately and efficiently extract text content from various images. It converts visual text information in images into machine-readable, editable, and searchable structured data.

Request Headers

  • Name
    Content-Type
    Type
    string
    Required
    Required
    Description
    The media type of the request body, which determines the format of the image field. Supports the following two values:
    • application/json:When using this type, the image field in the request body must be a Base64 encoded string.
    • multipart/form-data:When using this type, the image field in the request body must be uploaded as a file.
  • Name
    Authorization
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. The credential used for API authentication. For details, see Authentication.
  • Name
    x-api-key
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. Pass your API key directly. Note: Do not use both Authorization and x-api-key at the same time.

Request Body

  • Name
    model
    Type
    string
    Required
    Required
    Description
    The model ID used for optical character recognition. Please choose the appropriate model based on your needs. Currently supported models include image-ocr-lite and image-ocr-base.
  • Name
    image
    Type
    string | file
    Required
    Required
    Description
    The image source for text recognition. The format of this field must match the Content-Type value in the request header:
    • When Content-Type is application/json, the image field should be a Base64 encoded string.
    • When Content-Type is multipart/form-data, the image field should be an uploaded image file.

Response Body

  • Name
    text
    Type
    string
    Required
    Required
    Description
    All text content recognized by the model.

Request

POST
/beta/v1/images/ocr
curl 'https://api.easytransnote.com/beta/v1/images/ocr' \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $YOUR_API_KEY" \
  -d '{
    "model": "image-ocr-base",
    "image": "iVBORw0KGgoAAAANSUhEUgAAAA...SUVORK5CYII="
  }'

Response

{
  "text": "这里是描述出的所有文本内容..."
}

POST/beta/v1/images/describe

Image Describe

The Image Describe API is an advanced visual language model interface designed for in-depth image content analysis. It can identify objects, environments, actions, and even emotions in images, generating precise, rich, and human-language-compliant text descriptions. This API is dedicated to building a bridge between vision and language, giving your application the ability to truly "understand" images.

Request Headers

  • Name
    Content-Type
    Type
    string
    Required
    Required
    Description
    The media type of the request body, which determines the format of the image field. Supports the following two values:
    • application/json:When using this type, the image field in the request body must be a Base64 encoded string.
    • multipart/form-data:When using this type, the image field in the request body must be uploaded as a file.
  • Name
    Authorization
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. The credential used for API authentication. For details, see Authentication.
  • Name
    x-api-key
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. Pass your API key directly. Note: Do not use both Authorization and x-api-key at the same time.

Request Body

  • Name
    model
    Type
    string
    Required
    Required
    Description
    The model ID used for image content description. Please choose the appropriate model based on your needs. Currently supported models include image-describe-lite and image-describe-base.
  • Name
    image
    Type
    string | file
    Required
    Required
    Description
    The image source for text recognition. The format of this field must match the Content-Type value in the request header:
    • When Content-Type is application/json, the image field should be a Base64 encoded string.
    • When Content-Type is multipart/form-data, the image field should be an uploaded image file.

Response Body

  • Name
    text
    Type
    string
    Required
    Required
    Description
    All text content recognized by the model.

Request

POST
/beta/v1/images/describe
curl 'https://api.easytransnote.com/beta/v1/images/describe' \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $YOUR_API_KEY" \
  -d '{
    "model": "image-describe-base",
    "image": "iVBORw0KGgoAAAANSUhEUgAAAA...SUVORK5CYII="
  }'

Response

{
  "text": "这里是描述出的所有文本内容..."
}

Was this page helpful?