Messages
The Messages API is Anthropic's core interface for building complex, multi-turn conversations. It constructs around a messages array containing alternating user and assistant roles, enabling precise simulation and continuation of conversation history. This interface supports multimodal inputs such as text and images, and provides persistent, high-priority instructions through a top-level system parameter.
With an OriginRouter One subscription, choose the corresponding Coding API endpoint and make sure the model ID comes from the Supported Models list; with the pay-as-you-go API plan, choose the corresponding Beta API endpoint and make sure the model ID comes from the Model List; a plan and endpoint mismatch may make models unavailable or result in unexpected billing.
Create a Message
This endpoint creates a model message based on the content you provide.
Request Headers
- Name
Content-Type- Type
- string
- Required
- Required
- Description
- Value must be
application/json.
- Name
anthropic-version- Type
- string
- Required
- Required
- Description
- The Anthropic API version you want to use, currently only
2023-06-01is supported.
- Name
Authorization- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. Credentials for API authentication, see Authentication for details.
- Name
x-api-key- Type
- string
- Optional
- Optional
- Description
- Optional authentication method. Pass your API key directly. Note: Do not use both
Authorizationandx-api-keysimultaneously.
Request Body
- Name
model- Type
- string
- Required
- Required
- Description
- The model ID to use. Available values include claude-3-5-sonnet-20241022, claude-3-5-haiku-20241022, claude-3-opus-20240229, claude-3-sonnet-20240229, claude-3-haiku-20240307, etc.
- Name
messages- Type
- array
- Required
- Required
- Description
- Input messages.
When creating a new message, you need to specify previous conversation turns through the
messagesparameter, and the model will then generate the next message in the conversation. Consecutiveuserorassistantturns in the request will be merged into a single turn.Each input message must be an object containing
roleandcontent. You can specify a singleuserrole message or include multipleuserandassistantmessages.If the last message uses the
assistantrole, the response content will continue directly from that message's content. This can be used to constrain part of the model's response.Note: The maximum number of messages per request is 100,000.
- Name
max_tokens- Type
- integer
- Required
- Required
- Description
- The maximum number of tokens to generate before stopping. Minimum value is 1, and depending on the model, the maximum is typically 4096 or higher.Note: The model may stop generating before reaching this maximum value. This parameter only specifies the absolute maximum number of tokens that can be generated.
- Name
system- Type
- string
- Optional
- Optional
- Description
- System prompts are a way to provide context and instructions to the model, such as specifying a specific goal or role.System prompts are used to set the behavior and context of the model, taking priority over user messages in the messages array.
- Name
temperature- Type
- number
- Optional
- Optional
- Description
- The degree of randomness injected into the response. Default is 1.0, with a range of 0.0 to 1.0.Lower values (such as 0.0-0.3) make output more deterministic and consistent, while higher values (such as 0.7-1.0) make output more diverse and creative. The default value of 1.0 is suitable for general conversations.
- Name
top_k- Type
- integer
- Optional
- Optional
- Description
- Only sample from the top K options for each subsequent token. Must be a positive integer.Used to remove "long-tail" low-probability responses. Only recommended for advanced use cases. Usually you only need to use the temperature parameter.
- Name
top_p- Type
- number
- Optional
- Optional
- Description
- Uses nucleus sampling. Allowed range is 0.0 to 1.0. Used to limit tokens considered based on cumulative probability.In nucleus sampling, we calculate the cumulative distribution of all options for each subsequent token from highest to lowest probability, and truncate when reaching the specific probability specified by
top_p. You should adjust temperature ortop_p, but not both.
- Name
tools- Type
- array
- Optional
- Optional
- Description
- Definitions of tools the model may use.If you include tools in your API request, the model may return
tool_usecontent blocks indicating the model used these tools. You can run these tools using the tool inputs generated by the model, and optionally return the results to the model usingtool_resultcontent blocks.
- Name
tool_choice- Type
- object | string
- Optional
- Optional
- Description
- Controls which tools the model uses. Supported values:
auto(model decides),any(model must use a tool),{"type": "tool", "name": "tool_name"}(use specific tool),{"type": "none"}(don't use tools). Default isauto.
- Name
stream- Type
- boolean
- Optional
- Optional
- Description
- Whether to use Server-Sent Events (SSE) to stream responses incrementally. Default is
false.
- Name
stop_sequences- Type
- array<string>
- Optional
- Optional
- Description
- Custom text sequences that trigger the model to stop generation.Our models usually stop when they naturally complete their turn, which will cause the response stop reason
stop_reasonto beend_turn.If you want the model to stop when encountering custom text sequences, you can use thestop_sequencesparameter. If the model encounters one of these custom sequences, the response stop reasonstop_reasonwill bestop_sequence, and the response stop sequencestop_sequencewill contain the matched stop sequence.
- Name
metadata- Type
- object
- Optional
- Optional
- Description
- An object describing request metadata.
- Name
thinking- Type
- object
- Optional
- Optional
- Description
- Enable the model's extended thinking capability.When enabled, the response will include thinking content blocks showing Claude's thought process before giving the final answer. This feature requires at least 1,024 tokens and counts toward your
max_tokenslimit.
- Name
multimodal- Type
- object
- Optional
- Optional
- Description
- Multimodal adaptation configuration. Used to enable and control automatic conversion of non-text content (such as images, PDFs, videos, audio, etc.). For details, see Multimodal Support.
- Name
fallback- Type
- string
- Optional
- Optional
- Description
- Defines the fallback strategy when a request fails. This parameter only applies to the
/beta/v1/messagesendpoint. For details, see Model Fallback.
- Name
fallback_config- Type
- object
- Optional
- Optional
- Description
- Detailed configuration for custom fallback behavior. This parameter only applies to the
/beta/v1/messagesendpoint. For details, see Model Fallback.
- Name
thinking_block_strategy- Type
- string | null
- Optional
- Optional
- Description
- Controls the strategy the server uses to handle thinking content blocks in the
messagesparameter of the current request.For most cross-route requests, thethinking.signaturegenerated by different provider models are not interchangeable, and passing them directly will cause errors. By setting this strategy, you can flexibly control how thinking content is handled.Defaults topreservefor the/beta/v1/messagesendpoint and the user's subscription setting (typicallypreserve) for the/coding/v1/messagesendpoint. You can also control the default behavior of this feature in the Console.The currently supported strategy values are as follows:- preserve: Preserves all thinking inputs with valid signatures, and automatically removes improperly formatted or unsigned thinking inputs.
- discard: Discards all incoming thinking inputs. This mode is compatible with almost all models and routing protocols, but the output quality may be reduced.
- flatten: Converts all incoming thinking inputs into plain text inputs. While this mode can effectively improve output quality, it significantly reduces inference cache hit rates and substantially increases per-request latency.
- Name
use_built_in_tools- Type
- boolean
- Optional
- Optional
- Description
- Controls whether server built-in tools are enabled. Defaults to
falsefor the/beta/v1/messagesendpoint and the user's subscription setting (typicallytrue) for the/coding/v1/messagesendpoint.Most LLMs on model provider routes do not have the ability to automatically invoke external tools such as web search. When enabled, the server automatically injects external tool definitions such as web search, handles tool calls and result processing within a single request, and returns the complete response. You can also control the default behavior of this feature in the Console.
Response Body
- Name
id- Type
- string
- Optional
- Optional
- Description
- The unique identifier for this message, e.g.,
msg_01EcyWo6m4hyW8KHs2y2pei5.
- Name
type- Type
- string
- Optional
- Optional
- Description
- The object type, for this endpoint its value is always
message.
- Name
role- Type
- string
- Optional
- Optional
- Description
- The role of the message author. For response messages, this value is always
assistant.
- Name
content- Type
- array
- Optional
- Optional
- Description
- An array of content blocks that make up the message content. This block-based structure allows mixing different types of content in a single response.
- Name
model- Type
- string
- Optional
- Optional
- Description
- The name of the model that processed this request, e.g.,
gpt-5.
- Name
stop_reason- Type
- string | null
- Optional
- Optional
- Description
- The reason the model stopped generating content. This is a key field for controlling application flow. Possible values include:
- end_turn: Model reached a natural stopping point.
- max_tokens: The requested max_tokens limit was reached.
- stop_sequence: One of your custom stop_sequences was generated.
- Name
stop_sequence- Type
- string | null
- Optional
- Optional
- Description
- If generation was interrupted by a custom stop sequence, this field will contain that sequence.
- Name
usage- Type
- object
- Optional
- Optional
- Description
- Detailed billing and rate limit usage information, providing higher transparency than OpenAI.
Request
curl https://api.easytransnote.com/beta/v1/messages \
-H "x-api-key: $API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
--no-buffer \
-d '{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello, Claude!"}
]
}'
Response
{
"id": "msg_1a2b3c4d5e6f",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
],
"model": "claude-3-5-sonnet-20241022",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 10,
"output_tokens": 8,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}