Images
The Images API provides a comprehensive set of AI-powered image capabilities, including image generation, image editing, OCR text recognition, and image description. With these APIs, you can generate images from text, modify existing images, extract text from images, and even describe image content using natural language.
With an OriginRouter One subscription, choose the corresponding Coding API endpoint and make sure the model ID comes from the Supported Models list; with the pay-as-you-go API plan, choose the corresponding Beta API endpoint and make sure the model ID comes from the Model List; a plan and endpoint mismatch may make models unavailable or result in unexpected billing.
Image Generation
The Image Generation API is the core of text-to-image capabilities. Simply provide a detailed text description, and the model will create original images that match your description. This is ideal for building art generators, marketing poster designs, and product prototype visualizations.
Request Headers
- Name
Content-Type- Type
- string
- Required
- Required
- Description
- Value must be
application/json.
- Name
Authorization- Type
- string
- Required
- Required
- Description
- Optional authentication method. The credential used for API authentication. For details, see Authentication.
Request Body
- Name
model- Type
- string
- Required
- Required
- Description
- The model ID used for image generation. Please choose the appropriate model based on your needs. Currently supported models include
gpt-image-1andgemini-2.5-flash-image-preview.
- Name
prompt- Type
- string
- Required
- Required
- Description
- The text description of the desired image.
gpt-image-1has a maximum length of 32000 characters, anddall-e-3has a maximum of 4000 characters.
- Name
background- Type
- string
- Optional
- Optional
- Description
- Allows setting background transparency for the generated image. This parameter only applies to
gpt-image-1. Must be one oftransparent,opaque, orauto(default). When usingauto, the model will automatically determine the best background for the image.Iftransparentis selected, the output format must support transparency, so it should be set to png (default) or webp.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
moderation- Type
- string
- Optional
- Optional
- Description
- Controls the content moderation level of images generated by the model. Must be set to
low(less filtering restrictions) orauto(default).
- Name
n- Type
- string
- Optional
- Optional
- Description
- The number of images to generate. Must be between 1 and 10. For
dall-e-3andgemini-2.5-flash-image-preview, onlyn=1is supported.
- Name
output_compression- Type
- integer
- Optional
- Optional
- Description
- The compression level of the generated image (0-100%). This parameter only applies to the
gpt-image-1model with output format webp or jpeg. Default is 100.
- Name
output_format- Type
- string
- Optional
- Optional
- Description
- The format of the generated image returned. This parameter only applies to
gpt-image-1. Must be one of png (default), jpeg, or webp.
- Name
partial_images- Type
- string
- Optional
- Optional
- Description
- The number of partial images to generate. This parameter is used to return a streaming response with partial images. The value must be between 0 and 3. When set to 0, the response will be sent as a single image in one streaming event.Note: If the complete image is generated faster, the final image may be sent before all partial images are generated.
- Name
quality- Type
- string
- Optional
- Optional
- Description
- The quality of the generated image.
auto(默认值)将自动为给定模型选择最佳质量。gpt-image-1supportshigh,medium, andlowquality levels.dall-e-3supportshdandstandardquality levels.
gemini-2.5-flash-image-previewmodel.
- Name
response_format- Type
- string
- Optional
- Optional
- Description
- The return format of images generated by DALL-E 3. Must be one of
urlorb64_json. URLs are valid for only 60 minutes after image generation. This parameter does not apply togpt-image-1andgemini-2.5-flash-image-preview, which always return base64 encoded images.
- Name
size- Type
- string
- Optional
- Optional
- Description
- The size of the generated image. For
gpt-image-1, must be one of1024x1024,1536x1024(landscape),1024x1536(portrait), orauto(default); fordall-e-3, it is1024x1024,1792x1024, or1024x1792.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
stream- Type
- string
- Optional
- Optional
- Description
- Generate images in streaming mode. Default is
false. This parameter does not apply todall-e-3.
- Name
style- Type
- string
- Optional
- Optional
- Description
- The style of the generated image. This parameter only applies to the
dall-e-3model. Must choose one ofvividornatural. Choosing "vivid" will make the model tend to generate hyper-realistic and dramatic images; choosing "natural" will make the model generate more natural, less hyper-realistic images.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
user- Type
- string
- Optional
- Optional
- Description
- A unique identifier representing the end user, used for monitoring and detecting abuse.
Response Body
- Name
background- Type
- string
- Required
- Required
- Description
- The background parameter used for image generation. Value is
transparentoropaque.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
created- Type
- integer
- Required
- Required
- Description
- Unix timestamp (in seconds) when the image was created.
- Name
data- Type
- array
- Required
- Required
- Description
- List of generated images.
- Name
output_format- Type
- string
- Required
- Required
- Description
- Output format for image generation. Values are png, webp, or jpeg.
- Name
quality- Type
- string
- Required
- Required
- Description
- The quality of the generated image. Values are
low,medium, orhigh.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
size- Type
- string
- Required
- Required
- Description
- The size of the generated image. Values are
1024x1024,1024x1536, or1536x1024.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
usage- Type
- object
- Required
- Required
- Description
- Token usage information for image generation. This parameter is not supported for the
dall-e-3model.
Request
curl https://api.easytransnote.com/beta/v1/images/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "gpt-image-1",
"prompt": "A cute baby sea otter",
"n": 1,
"size": "1024x1024"
}'
Response
{
"created": 1757908389,
"data": [
{
"b64_json": "iVBORw0KGgoAAAANSUhEUgA"
}
],
"usage": {
"input_tokens": 75,
"input_tokens_details": {
"image_tokens": 0,
"text_tokens": 75
},
"output_tokens": 1024,
"total_tokens": 1099
}
}
Image Edit
The Image Edit API provides powerful image restoration and modification capabilities. Unlike generating from scratch, this API allows you to upload an original image and precisely modify specific areas through text instructions and optional masks. This API is widely used in photo restoration, product image refinement, creative composition, and other scenarios.
Request Headers
- Name
Content-Type- Type
- string
- Required
- Required
- Description
- Value must be
multipart/form-data.
- Name
Authorization- Type
- string
- Required
- Required
- Description
- Optional authentication method. The credential used for API authentication. For details, see Authentication.
Request Body
- Name
model- Type
- string
- Required
- Required
- Description
- The model ID used for image generation. Please choose the appropriate model based on your needs. Currently supported models include
gpt-image-1andgemini-2.5-flash-image-preview.
- Name
prompt- Type
- string
- Required
- Required
- Description
- The text description of the desired image.
gpt-image-1has a maximum length of 32000 characters.
- Name
image- Type
- file | array
- Required
- Required
- Description
- The image file to be edited. Must be a supported image format or an array of images.For
gpt-image-1, each image should be in png, webp, or jpg format and less than 50MB. Up to 16 images can be provided.
- Name
background- Type
- string
- Optional
- Optional
- Description
- Allows setting background transparency for the generated image. This parameter only applies to
gpt-image-1. Must be one oftransparent,opaque, orauto(default). When usingauto, the model will automatically determine the best background for the image.Iftransparentis selected, the output format must support transparency, so it should be set to png (default) or webp.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
input_fidelity- Type
- string
- Optional
- Optional
- Description
- Controls how much effort the model puts into matching the style and features (especially facial features) of the input image. This parameter only applies to the
gpt-image-1model. Supportshighandlowsettings, default islow.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
mask- Type
- file
- Optional
- Optional
- Description
- An additional image where completely transparent areas (e.g., where the Alpha channel value is zero) indicate the areas of the image that need to be edited. If multiple images are provided, the mask will be applied to the first image. Must be a valid PNG file, less than 4MB, and the same dimensions as the target image.Note: For model compatibility,
gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and additional mask images, and cannot guarantee the final image has this effect.
- Name
n- Type
- string
- Optional
- Optional
- Description
- The number of images to generate. Must be between 1 and 10. For
dall-e-3andgemini-2.5-flash-image-preview, onlyn=1is supported.
- Name
output_compression- Type
- integer
- Optional
- Optional
- Description
- The compression level of the generated image (0-100%). This parameter only applies to the
gpt-image-1model with output format webp or jpeg. Default is 100.
- Name
output_format- Type
- string
- Optional
- Optional
- Description
- The format of the generated image returned. This parameter only applies to
gpt-image-1. Must be one of png (default), jpeg, or webp.
- Name
partial_images- Type
- string
- Optional
- Optional
- Description
- The number of partial images to generate. This parameter is used to return a streaming response with partial images. The value must be between 0 and 3. When set to 0, the response will be sent as a single image in one streaming event.Note: If the complete image is generated faster, the final image may be sent before all partial images are generated.
- Name
quality- Type
- string
- Optional
- Optional
- Description
- The quality of the generated image.
auto(默认值)将自动为给定模型选择最佳质量。gpt-image-1supportshigh,medium, andlowquality levels.dall-e-3supportshdandstandardquality levels.
gemini-2.5-flash-image-previewmodel.
- Name
response_format- Type
- string
- Optional
- Optional
- Description
- The return format of images generated by DALL-E 3. Must be one of
urlorb64_json. URLs are valid for only 60 minutes after image generation. This parameter does not apply togpt-image-1andgemini-2.5-flash-image-preview, which always return base64 encoded images.
- Name
size- Type
- string
- Optional
- Optional
- Description
- The size of the generated image. For
gpt-image-1, must be one of1024x1024,1536x1024(landscape),1024x1536(portrait), orauto(default); fordall-e-3, it is1024x1024,1792x1024, or1024x1792.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, but the corresponding instruction is sent to the model via prompt and cannot guarantee the final image has this effect.
- Name
stream- Type
- string
- Optional
- Optional
- Description
- Generate images in streaming mode. Default is
false.
- Name
user- Type
- string
- Optional
- Optional
- Description
- A unique identifier representing the end user, used for monitoring and detecting abuse.
Response Body
- Name
background- Type
- string
- Required
- Required
- Description
- The background parameter used for image generation. Value is
transparentoropaque.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
created- Type
- integer
- Required
- Required
- Description
- Unix timestamp (in seconds) when the image was created.
- Name
data- Type
- array
- Required
- Required
- Description
- List of generated images.
- Name
output_format- Type
- string
- Required
- Required
- Description
- Output format for image generation. Values are png, webp, or jpeg.
- Name
quality- Type
- string
- Required
- Required
- Description
- The quality of the generated image. Values are
low,medium, orhigh.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
size- Type
- string
- Required
- Required
- Description
- The size of the generated image. Values are
1024x1024,1024x1536, or1536x1024.Note: For model compatibility,gemini-2.5-flash-image-previewalso supports this parameter, and the corresponding value is alwaysauto.
- Name
usage- Type
- object
- Required
- Required
- Description
- Token usage information for image generation. This parameter is not supported for the
dall-e-3model.
Request
curl -s -D >(grep -i x-request-id >&2) \
-o >(jq -r '.data[0].b64_json' | base64 --decode > gift-basket.png) \
-X POST "https://api.easytransnote.com/beta/v1/images/edits" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-F "model=gpt-image-1" \
-F "image[][email protected]" \
-F "image[][email protected]" \
-F "image[][email protected]" \
-F "image[][email protected]" \
-F 'prompt=Create a lovely gift basket with these four items in it'
Response
{
"created": 1757908389,
"data": [
{
"b64_json": "iVBORw0KGgoAAAANSUhEUgA"
}
],
"usage": {
"input_tokens": 75,
"input_tokens_details": {
"image_tokens": 0,
"text_tokens": 75
},
"output_tokens": 1024,
"total_tokens": 1099
}
}
Image OCR
The Image OCR API is a powerful interface designed to accurately and efficiently extract text content from various images. It converts visual text information in images into machine-readable, editable, and searchable structured data.
Request Headers
- Name
Content-Type- Type
- string
- Required
- Required
- Description
- The media type of the request body, which determines the format of the image field. Supports the following two values:
application/json:When using this type, theimagefield in the request body must be a Base64 encoded string.multipart/form-data:When using this type, theimagefield in the request body must be uploaded as a file.
- Name
Authorization- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. The credential used for API authentication. For details, see Authentication.
- Name
x-api-key- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. Pass your API key directly. Note: Do not use both
Authorizationandx-api-keyat the same time.
Request Body
- Name
model- Type
- string
- Required
- Required
- Description
- The model ID used for optical character recognition. Please choose the appropriate model based on your needs. Currently supported models include
image-ocr-liteandimage-ocr-base.
- Name
image- Type
- string | file
- Required
- Required
- Description
- The image source for text recognition. The format of this field must match the Content-Type value in the request header:
- When
Content-Typeisapplication/json, theimagefield should be a Base64 encoded string. - When
Content-Typeismultipart/form-data, theimagefield should be an uploaded image file.
- When
Response Body
- Name
text- Type
- string
- Required
- Required
- Description
- All text content recognized by the model.
Request
curl 'https://api.easytransnote.com/beta/v1/images/ocr' \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "image-ocr-base",
"image": "iVBORw0KGgoAAAANSUhEUgAAAA...SUVORK5CYII="
}'
Response
{
"text": "这里是描述出的所有文本内容..."
}
Image Describe
The Image Describe API is an advanced visual language model interface designed for in-depth image content analysis. It can identify objects, environments, actions, and even emotions in images, generating precise, rich, and human-language-compliant text descriptions. This API is dedicated to building a bridge between vision and language, giving your application the ability to truly "understand" images.
Request Headers
- Name
Content-Type- Type
- string
- Required
- Required
- Description
- The media type of the request body, which determines the format of the image field. Supports the following two values:
application/json:When using this type, theimagefield in the request body must be a Base64 encoded string.multipart/form-data:When using this type, theimagefield in the request body must be uploaded as a file.
- Name
Authorization- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. The credential used for API authentication. For details, see Authentication.
- Name
x-api-key- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. Pass your API key directly. Note: Do not use both
Authorizationandx-api-keyat the same time.
Request Body
- Name
model- Type
- string
- Required
- Required
- Description
- The model ID used for image content description. Please choose the appropriate model based on your needs. Currently supported models include
image-describe-liteandimage-describe-base.
- Name
image- Type
- string | file
- Required
- Required
- Description
- The image source for text recognition. The format of this field must match the Content-Type value in the request header:
- When
Content-Typeisapplication/json, theimagefield should be a Base64 encoded string. - When
Content-Typeismultipart/form-data, theimagefield should be an uploaded image file.
- When
Response Body
- Name
text- Type
- string
- Required
- Required
- Description
- All text content recognized by the model.
Request
curl 'https://api.easytransnote.com/beta/v1/images/describe' \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "image-describe-base",
"image": "iVBORw0KGgoAAAANSUhEUgAAAA...SUVORK5CYII="
}'
Response
{
"text": "这里是描述出的所有文本内容..."
}