API ResourcesMessages

Messages

The Messages API is Anthropic's core interface for building complex, multi-turn conversations. It constructs around a messages array containing alternating user and assistant roles, enabling precise simulation and continuation of conversation history. This interface supports multimodal inputs such as text and images, and provides persistent, high-priority instructions through a top-level system parameter.


POST/beta/v1/messages

Create a Message

This endpoint creates a model message based on the content you provide.

Request Headers

  • Name
    Content-Type
    Type
    string
    Required
    Required
    Description
    Value must be application/json.
  • Name
    anthropic-version
    Type
    string
    Required
    Required
    Description
    The Anthropic API version you want to use, currently only 2023-06-01 is supported.
  • Name
    Authorization
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. Credentials for API authentication, see Authentication for details.
  • Name
    x-api-key
    Type
    string
    Optional
    Optional
    Description
    Optional authentication method. Pass your API key directly. Note: Do not use both Authorization and x-api-key simultaneously.

Request Body

  • Name
    model
    Type
    string
    Required
    Required
    Description
    The model ID to use. Available values include claude-3-5-sonnet-20241022, claude-3-5-haiku-20241022, claude-3-opus-20240229, claude-3-sonnet-20240229, claude-3-haiku-20240307, etc.
  • Name
    messages
    Type
    array
    Required
    Required
    Description
    Input messages.

    When creating a new message, you need to specify previous conversation turns through the messages parameter, and the model will then generate the next message in the conversation. Consecutive user or assistant turns in the request will be merged into a single turn.

    Each input message must be an object containing role and content. You can specify a single user role message or include multiple user and assistant messages.

    If the last message uses the assistant role, the response content will continue directly from that message's content. This can be used to constrain part of the model's response.

    Note: The maximum number of messages per request is 100,000.

  • Name
    max_tokens
    Type
    integer
    Required
    Required
    Description
    The maximum number of tokens to generate before stopping. Minimum value is 1, and depending on the model, the maximum is typically 4096 or higher.Note: The model may stop generating before reaching this maximum value. This parameter only specifies the absolute maximum number of tokens that can be generated.
  • Name
    system
    Type
    string
    Optional
    Optional
    Description
    System prompts are a way to provide context and instructions to the model, such as specifying a specific goal or role.System prompts are used to set the behavior and context of the model, taking priority over user messages in the messages array.
  • Name
    temperature
    Type
    number
    Optional
    Optional
    Description
    The degree of randomness injected into the response. Default is 1.0, with a range of 0.0 to 1.0.Lower values (such as 0.0-0.3) make output more deterministic and consistent, while higher values (such as 0.7-1.0) make output more diverse and creative. The default value of 1.0 is suitable for general conversations.
  • Name
    top_k
    Type
    integer
    Optional
    Optional
    Description
    Only sample from the top K options for each subsequent token. Must be a positive integer.Used to remove "long-tail" low-probability responses. Only recommended for advanced use cases. Usually you only need to use the temperature parameter.
  • Name
    top_p
    Type
    number
    Optional
    Optional
    Description
    Uses nucleus sampling. Allowed range is 0.0 to 1.0. Used to limit tokens considered based on cumulative probability.In nucleus sampling, we calculate the cumulative distribution of all options for each subsequent token from highest to lowest probability, and truncate when reaching the specific probability specified by top_p. You should adjust temperature or top_p, but not both.
  • Name
    tools
    Type
    array
    Optional
    Optional
    Description
    Definitions of tools the model may use.If you include tools in your API request, the model may return tool_use content blocks indicating the model used these tools. You can run these tools using the tool inputs generated by the model, and optionally return the results to the model using tool_result content blocks.
  • Name
    tool_choice
    Type
    object | string
    Optional
    Optional
    Description
    Controls which tools the model uses. Supported values: auto (model decides), any (model must use a tool), {"type": "tool", "name": "tool_name"} (use specific tool), {"type": "none"} (don't use tools). Default is auto.
  • Name
    stream
    Type
    boolean
    Optional
    Optional
    Description
    Whether to use Server-Sent Events (SSE) to stream responses incrementally. Default is false.
  • Name
    stop_sequences
    Type
    array<string>
    Optional
    Optional
    Description
    Custom text sequences that trigger the model to stop generation.Our models usually stop when they naturally complete their turn, which will cause the response stop reason stop_reason to be end_turn.If you want the model to stop when encountering custom text sequences, you can use the stop_sequences parameter. If the model encounters one of these custom sequences, the response stop reason stop_reason will be stop_sequence, and the response stop sequence stop_sequence will contain the matched stop sequence.
  • Name
    metadata
    Type
    object
    Optional
    Optional
    Description
    An object describing request metadata.
  • Name
    thinking
    Type
    object
    Optional
    Optional
    Description
    Enable the model's extended thinking capability.When enabled, the response will include thinking content blocks showing Claude's thought process before giving the final answer. This feature requires at least 1,024 tokens and counts toward your max_tokens limit.
  • Name
    multimodal
    Type
    object
    Optional
    Optional
    Description
    Multimodal adaptation configuration. Used to enable and control automatic conversion of non-text content (such as images, PDFs, videos, audio, etc.). For details, see Multimodal Support.
  • Name
    fallback
    Type
    string
    Optional
    Optional
    Description
    Defines the fallback strategy when a request fails. This parameter only applies to the /beta/v1/messages endpoint. For details, see Model Fallback.
  • Name
    fallback_config
    Type
    object
    Optional
    Optional
    Description
    Detailed configuration for custom fallback behavior. This parameter only applies to the /beta/v1/messages endpoint. For details, see Model Fallback.
  • Name
    thinking_block_strategy
    Type
    string | null
    Optional
    Optional
    Description
    Controls the strategy the server uses to handle thinking content blocks in the messages parameter of the current request.For most cross-route requests, the thinking.signature generated by different provider models are not interchangeable, and passing them directly will cause errors. By setting this strategy, you can flexibly control how thinking content is handled.Defaults to preserve for the /beta/v1/messages endpoint and the user's subscription setting (typically preserve) for the /coding/v1/messages endpoint. You can also control the default behavior of this feature in the Console.The currently supported strategy values are as follows:
    • preserve: Preserves all thinking inputs with valid signatures, and automatically removes improperly formatted or unsigned thinking inputs.
    • discard: Discards all incoming thinking inputs. This mode is compatible with almost all models and routing protocols, but the output quality may be reduced.
    • flatten: Converts all incoming thinking inputs into plain text inputs. While this mode can effectively improve output quality, it significantly reduces inference cache hit rates and substantially increases per-request latency.
  • Name
    use_built_in_tools
    Type
    boolean
    Optional
    Optional
    Description
    Controls whether server built-in tools are enabled. Defaults to false for the /beta/v1/messages endpoint and the user's subscription setting (typically true) for the /coding/v1/messages endpoint.Most LLMs on model provider routes do not have the ability to automatically invoke external tools such as web search. When enabled, the server automatically injects external tool definitions such as web search, handles tool calls and result processing within a single request, and returns the complete response. You can also control the default behavior of this feature in the Console.

Response Body

  • Name
    id
    Type
    string
    Optional
    Optional
    Description
    The unique identifier for this message, e.g., msg_01EcyWo6m4hyW8KHs2y2pei5.
  • Name
    type
    Type
    string
    Optional
    Optional
    Description
    The object type, for this endpoint its value is always message.
  • Name
    role
    Type
    string
    Optional
    Optional
    Description
    The role of the message author. For response messages, this value is always assistant.
  • Name
    content
    Type
    array
    Optional
    Optional
    Description
    An array of content blocks that make up the message content. This block-based structure allows mixing different types of content in a single response.
  • Name
    model
    Type
    string
    Optional
    Optional
    Description
    The name of the model that processed this request, e.g., gpt-5.
  • Name
    stop_reason
    Type
    string | null
    Optional
    Optional
    Description
    The reason the model stopped generating content. This is a key field for controlling application flow. Possible values include:
    • end_turn: Model reached a natural stopping point.
    • max_tokens: The requested max_tokens limit was reached.
    • stop_sequence: One of your custom stop_sequences was generated.
  • Name
    stop_sequence
    Type
    string | null
    Optional
    Optional
    Description
    If generation was interrupted by a custom stop sequence, this field will contain that sequence.
  • Name
    usage
    Type
    object
    Optional
    Optional
    Description
    Detailed billing and rate limit usage information, providing higher transparency than OpenAI.

Request

POST
/beta/v1/messages
curl https://api.easytransnote.com/beta/v1/messages \
  -H "x-api-key: $API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  --no-buffer \
  -d '{
    "model": "claude-3-5-sonnet-20241022",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Hello, Claude!"}
    ]
  }'

Response

{
  "id": "msg_1a2b3c4d5e6f",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "Hello! How can I help you today?"
    }
  ],
  "model": "claude-3-5-sonnet-20241022",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 10,
    "output_tokens": 8,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

Was this page helpful?