Experimental Features
This page lists the experimental features provided by OriginRouter. These features are currently in active development or testing stages, aimed at collecting user feedback and validating new possibilities.
After January 1, 2026, using experimental features may require your primary account to have some additional permissions. You can apply for access to experimental features here.
Please note that before the official release, experimental features may be subject to their own data strategies and privacy policies. These specific policies may supplement or modify our main privacy policy and will take precedence. We strongly recommend that you review the detailed instructions and policy pages attached to each feature before use.
If you encounter any issues related to experimental features, please let us know here.
Multimodal Support
OriginRouter's multimodal adaptation feature addresses a common pain point: many powerful language models do not natively support all media formats. This feature acts as an intelligent adaptation layer in the background, automatically handling incompatible media types, allowing you to interact with all models through a unified API.
- Image Support: When you send a message containing
image_urlto a model that does not support images, the system automatically calls a high-performance vision model to parse the image, converting it into structured text descriptions, and seamlessly adds this description as context to the target model request. - Document Support: For models that do not support complex documents like PDF and DOCX, the system automatically selects an adaptation strategy based on the model's capabilities. For models with multimodal capabilities, documents are converted to images to preserve the original content layout as much as possible; for pure text models, a high-quality local OCR and document parsing service is enabled to extract content. For
text/*type files, they are uniformly converted to standard text blocks.
You can fine-tune the parsing strategy for various media formats by passing the multimodal parameter in the API request body. The system automatically determines this based on the target model's native capabilities. If the target model (such as a model with built-in vision capabilities) already natively supports a certain type of media, the system will automatically ignore your adaptation requirements and directly pass the native file to the model to ensure the best results and lowest latency.
Request Body
Request Body
- Name
multimodal- Type
- object
- Optional
- Optional
- Description
- Detailed configuration object for multimodal adaptation. If this parameter is not specified, image and PDF adaptation will be enabled by default.
- Name
enabled- Type
- boolean
- Optional
- Optional
- Description
- Global feature toggle. When set to
false, all multimodal adaptation conversions will be completely turned off. Default istrue.
- Name
image- Type
- object
- Optional
- Optional
- Description
- Image adaptation configuration.
- Name
pdf- Type
- object
- Optional
- Optional
- Description
- PDF document adaptation configuration.
- Name
video- Type
- object
- Optional
- Optional
- Description
- Video adaptation configuration.
- Name
audio- Type
- object
- Optional
- Optional
- Description
- Audio adaptation configuration.
Request
curl https://api.easytransnote.com/beta/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "gemini-2.5-pro",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "请总结这张图片的内容以及附件的 PDF。"
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/sample.jpg"
}
}
]
}
],
"multimodal": {
"enabled": true,
"image": {
"enabled": true,
"model": "image-describe-base"
},
"pdf": {
"enabled": true
}
}
}'
Parameter Validation and Limits: When global enabled is true, at least one of the four sub-items (image, pdf, video, audio) must remain enabled, otherwise the API will return an invalid_parameter_value (400) error.
Since version beta-0-6-3-260226, the multimodal support feature has been opened in the testing stage and can currently only be used through the test endpoint /beta/v1/chat/completions. To confirm the current service version, please refer to Get Version.
Cross-Platform File System
Managing and transferring files between different AI providers (such as OpenAI, Anthropic, Google) often faces issues like complex processes, high costs, and repeated uploads. OriginRouter provides a unified file API that completely abstracts and simplifies cross-platform file processing logic.
We are compatible with OpenAI and Anthropic's standard file interfaces (such as /beta/v1/files, /beta/v1/files/{file_id}, and /beta/v1/files/{file_id}/content). The system will automatically identify the protocol format and adapt based on your request headers.
With OriginRouter's cross-platform file system, you can easily achieve file reuse across vendor model calls, reduce network overhead from large file reads, and avoid repeated uploads or file management across multiple platforms. OriginRouter will internally handle all file storage, distribution, and compatibility processing.
Privacy and Security Notice: All files you upload are processed with RSA symmetric encryption before entering our storage system, and are only decrypted in a controlled environment when needed for access. According to our privacy policy, we will not proactively review the content of files you upload. However, please note that the privacy policies and terms of service of our underlying storage service provider (Alibaba Cloud OSS) shall prevail according to its official documentation.
Since version beta-0-6-1-0808, the file interface has migrated from the testing stage endpoints /beta/v1/files, /beta/v1/files/{file_id}, and /beta/v1/files/{file_id}/content to the official interface paths. To view the specific API documentation, please refer to Files. To confirm the current service version, please refer to Get Version.
Model Fallback
In production environments, the stability of API calls is crucial. Large model calls may sometimes fail due to upstream provider overload, rate limits, or content moderation. The model fallback feature allows you to pre-configure alternative model routes. When the primary model call fails, the system will automatically switch and attempt to call an alternative model, significantly improving service stability and reliability.
You can specify the fallback strategy by passing custom fallback and fallback_config parameters in the API request body.
Request Body
Request Body
- Name
fallback- Type
- string
- Optional
- Optional
- Description
- Defines the fallback strategy when a request fails. Default is
auto.disabled: Disables the fallback mechanism.auto: Automatic mode, ignores thefallback_configparameter, and is handled by the system automatically.enabled: Enables fallback, using user-definedfallback_configconfiguration.strict: Strict mode, enables fallback and usesfallback_config. In this mode, the outermostmodelparameter will be ignored, and it will start trying from step 1 in the configuration.
- Name
fallback_config- Type
- object
- Optional
- Optional
- Description
- Detailed configuration for custom fallback mechanism. Only takes effect when
fallbackis set toenabledorstrict.- Name
strategy- Type
- object
- Required
- Required
- Description
- Global execution strategy, defining how to schedule the various steps in
steps.
- Name
global_timeout_ms- Type
- integer
- Optional
- Optional
- Description
- Global circuit breaker timeout (in milliseconds). The entire fallback chain (including all steps and retries) must complete within this time, otherwise the task will be forcefully terminated. For example, 180000 represents 3 minutes.
- Name
max_total_retries- Type
- integer
- Optional
- Optional
- Description
- The upper limit of the sum of all step retry counts allowed in the entire fallback chain. Used to prevent resource abuse due to configuration errors.
- Name
steps- Type
- array
- Required
- Required
- Description
- Defines the list of specific execution steps.
Request
curl https://api.easytransnote.com/beta/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "gemini-2.5-pro",
"messages": [
{
"role": "user",
"content": "2020年世界杯冠军是谁?"
}
],
"temperature": 0.7,
"max_tokens": 150,
"fallback": "enabled",
"fallback_config": {
"strategy": {
"mode": "sequential"
},
"global_timeout_ms": 180000,
"max_total_retries": 5,
"steps": [
{
"step_id": "step1_claude",
"model": "claude-sonnet-4-5-20250929",
"timeout_ms": 30000,
"retry_policy": {
"max_retries": 2,
"backoff_factor": 2,
"jitter_ms": 250
},
"fail_on": {
"status_code": [400, 401, 403],
"error_contains": ["invalid_request_error", "unauthenticated"],
},
"parameters": {
"temperature": 1
}
}
]
}
}'
Since version beta-0-6-2-1118, the model fallback feature has been opened in the testing stage and can currently only be used through the test endpoint /beta/v1/chat/completions. To confirm the current service version, please refer to Get Version.
Adaptive Memory
Traditional large language model APIs usually operate in a stateless manner, which requires developers to explicitly maintain conversation context, or accept that the model forgets user preferences and historical information when context is limited.
OriginRouter's Adaptive Memory Engine introduces a persistent cognitive mechanism at the /chat/completions interface layer for structured processing and management of conversation information. While supporting cross-session semantic continuity, this mechanism efficiently organizes and reuses conversation context, allowing the model to significantly reduce token usage costs in long conversation scenarios while maintaining reasoning consistency.
This feature includes two core dimensions of memory capability: Session Context Cache and Dynamic User Profiling.
-
Session Context Cache: OriginRouter's context cache is based on a self-designed message grouping algorithm that can identify complete interaction units in conversations and manage them in a structured way. In long conversation scenarios, this mechanism can significantly reduce model inference costs while maintaining semantic coherence.
- Long Conversation Optimization and Token Savings: The system automatically prunes redundant information, providing only logically complete interaction units as model input, thereby reducing token consumption per API call and improving inference efficiency.
- Key Settings Lock: In multi-turn interactions, the system keeps initial system instructions and key constraint conditions unchanged, preventing the sliding window truncation from causing the model to forget identity or task constraints, thus ensuring behavioral consistency in long-cycle interactions.
- Cross-Model and Flow Adaptation: Historical conversation data and context management logic can be seamlessly applied to different underlying large models without refactoring existing code, and without requiring separate adaptation for knowledge bases or third-party business processes, reducing technical migration and maintenance costs.
-
Dynamic User Profiling (Holographic User Profiling): OriginRouter can transform each interaction with AI into a systematic user profile, achieving long-term management of user preferences and behavioral patterns. The system extracts and structures key information from historical conversations, allowing the model to continue context and adapt to user characteristics in subsequent sessions.
- Cross-Session Long-Term Memory: Breaks through the limitations of traditional stateless models that "forget after chatting." Even in new sessions months later, the model can still recognize user historical preferences and directly provide suggestions based on past information.
- Intelligent Adaptation and Personalization: The model can automatically adjust response strategies based on user characteristics. For example, providing more explanatory information for beginners, and more direct technical implementation solutions for experienced developers.
- Zero-Intervention Personalization: Users don't need to manually configure or repeatedly input background information; the system automatically maintains consistency of user identity, preferences, and behavioral patterns, achieving global memory synchronization.
Integration Guide:
When enabling the memory feature, you need to provide memory_id in the request body or extension parameters. This parameter is used to identify and manage specific user or session memories. The system will persistently store and retrieve related information based on this ID, thereby supporting cross-session semantic continuity. This feature is currently in Beta testing stage, and API behavior and data structures may change with version updates. Developers should ensure proper exception handling during integration and refer to the latest specifications in the API documentation to ensure compatibility.
Request Body
Request Body
- Name
memory_id- Type
- string
- Required
- Required
- Description
- Unique identifier for the memory entity. Users need to generate it themselves and ensure its uniqueness within the user's namespace.After enabling this parameter, the server will fully manage conversation history. In subsequent
/chat/completionscalls, the system supports full upload (complete messages) or incremental upload (only the latest messages), both have the same effect, and the system will automatically handle context concatenation.Currently, this feature is free during the Beta period. Quota limits for different subscription levels are as follows:- Personal Free: 10
- Personal Pro: 50
- Team Pro / Enterprise: 100
Lifecycle Management: If a memory entity has no read/write interaction for more than 30 days, that memory entity and its associated data will be automatically destroyed.
- Name
memory_config- Type
- object
- Optional
- Optional
- Description
- Configuration object for fine-grained control of the memory system behavior.
- Name
memory_mode- Type
- read_write | read_only
- Optional
- Optional
- Description
- Sets the memory interaction mode. Default is
read_write.- read_write: Utilizes history in real-time to assist generation and updates the memory store based on new conversations. Uses a strong consistency lock (concurrency limit of 1), which is held until the current request fully completes. Messages are vectorized and summarized in a background async queue (typically within 10 minutes).
- read_only: Only retrieves memory to assist generation, without writing or updating. Provides higher throughput (concurrency limit of 5).
Data Consistency Verification: In read-write mode, the system performs strict hash verification on the
messagespassed by the client to ensure strict consistency between the client context and server storage.Special Role Description:
systemanddevelopermessages do not participate in the memory system's recording, processing, and hash verification. However, to ensure the coherence of model settings, users must include completesystemanddevelopermessages in each request. The system will directly apply these instructions without including them in historical consistency comparison.For conversation history (i.e.,
userandassistantmessages), the first uploaded list is treated as the baseline, and subsequent requests must follow the "Append Only" principle.Verification Example: Assuming the conversation history stored on the server is
[A, B]:- ✅
[System, A, B, C, D]: Passed (System ignored, A, B matched, appended C, D) - ✅
[System, B, C, D]: Passed (System ignored, sliding window matched, appended C, D) - ✅
[System, C, D]: Passed (System ignored, pure incremental mode, appended C, D) - ❌
[System, A, C, D]: Rejected (B missing in the middle, considered historically discontinuous)
Note: It is strictly prohibited to modify stored historical message content (User/Assistant). Any hash mismatch caused by modification will result in request rejection.
- Name
history_window- Type
- integer
- Optional
- Optional
- Description
- Short-term memory window size (unit: conversation turns). Default is 10, minimum is 5.The latest
history_windowgroups of conversations within this window will retain Raw Text, not compressed, and be sent to the large model in full to ensure the accuracy of the recency effect.In read-write mode, if the archiving of historical records outside the window is not completed, the system will block the request until consistency synchronization is completed.1. When the current cumulative turns are less than the set value, the memory compression logic is not triggered. 2. Immutable parameter: This parameter only takes effect during memory entity initialization. Passing this parameter in subsequent requests will be ignored.
- Name
summary_window- Type
- integer
- Optional
- Optional
- Description
- Summary trigger step size (unit: conversation turns). Default is 5, minimum is 5.Whenever the newly generated conversation turns reach
summary_windowgroups and slide out of thehistory_windowprotection zone, the system automatically triggers an async task to perform semantic compression and summarization on this part of the history.1. The summarization algorithm will never process active messages within thehistory_windowprotection period. 2. Immutable parameter: This parameter only takes effect during memory entity initialization. Passing this parameter in subsequent requests will be ignored.
- Name
personal_memory- Type
- boolean
- Optional
- Optional
- Description
- Whether to enable personal profile memory. Default is
false.This feature is independent of specificmemory_idand is used to maintain user-level preferences and settings. When enabled, the system automatically extracts user characteristics (such as profession, language style, special requirements) from conversations and shares them across different sessions to provide more personalized responses.Personal memory updates are also controlled bymemory_mode: in read-write mode, profiles are continuously corrected; in read-only mode, profiles are only read for assisted generation.
Request
curl https://api.easytransnote.com/beta/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $YOUR_API_KEY" \
-d '{
"model": "gemini-2.5-pro",
"messages": [
{
"role": "user",
"content": "记得我之前提到的项目代号吗?请帮我生成一份周报大纲。"
}
],
"temperature": 0.7,
"memory_id": "mem_user_123456_project_beta",
"memory_config": {
"memory_mode": "read_write",
"history_window": 10,
"summary_window": 5,
"personal_memory": true
}
}'
Since version beta-0-6-2-260110, the memory feature has been opened in the testing stage and can currently only be used through the test endpoint /beta/v1/chat/completions. To confirm the current service version, please refer to Get Version.