Avatar
novita/ling-30-flash-santeProtected
Circuit
Versions
API
RUNNING
The circuit is running and is healthy.
3Total Runs
6,175Avg Response Time (ms)
206Avg Tokens per Second
0.00%Error Rate
Control Panel
Login to ModelWorksLogin to ModelWorks to use this circuit. To get started, get an API key.
Explore endpointsModelWorks allows apps to access LLMs and MCP servers worldwide easily in any datacenter, securely and privately.
SecretsOverride default secrets and set up your own for better billing and security.
Circuit Card

Ling 3.0 Flash Sante

Novita AI (Ling 3.0 Flash Sante)

Overview

Ling 3.0 Flash Sante is a health and medicine-enhanced MoE model from InclusionAI with 124B total parameters and 5.1B activated per token. Optimized for medical knowledge reasoning, clinical safety, and evidence-based retrieval. Supports function calling with a 256K token context window — served via Novita AI serverless inference. This circuit provides an OpenAI-compatible chat completion endpoint backed by the inclusionai/ling-3.0-flash-sante model.


Inputs

1. Messages (messages)

  • Description: A list of messages comprising the conversation so far.
  • Type: array
  • Required: Yes

2. Max Tokens (max_tokens)

  • Description: The maximum number of tokens to generate in the completion. Ling 3.0 Flash Sante supports up to 32768.
  • Type: number
  • Default Value: 32768
  • Required: No

3. Temperature (temperature)

  • Description: Sampling temperature, controls randomness. Range [0.0, 2.0]. Lower values are more deterministic.
  • Type: number
  • Default Value: 1
  • Required: No

4. Top-p (top_p)

  • Description: Nucleus sampling. Range [0.01, 1.0]. Recommended to use either this or temperature, not both.
  • Type: number
  • Default Value: 0.95
  • Required: No

5. Tools (tools)

  • Description: A list of tools the model may call. Supports function-type tools.
  • Type: array
  • Default Value: []
  • Required: No

6. Tool Choice (tool_choice)

  • Description: Controls how the model selects a tool. auto or none.
  • Type: string
  • Default Value: auto
  • Required: No

Outputs

1. Chat Completion Output (chat_completion_output)

  • Description: The generated chat completion text — the model's final answer.
  • Value: Accumulated from choices[0].delta.content stream chunks.

2. Reasoning Output (reasoning_output)

  • Description: The chain-of-thought reasoning content — the model's thinking process, separate from the final answer.
  • Value: Accumulated from choices[0].delta.reasoning_content stream chunks.

3. Chat Completion Tool Calls (chat_completion_tool_calls)

  • Description: Tool calls made by the model (if any).
  • Value: Accumulated from choices[0].delta.tool_calls stream chunks.

Try It

You can try this circuit directly in the Playground — no setup required.


Notes

  1. Text-Only Input: Ling 3.0 Flash Sante does not support image, video, or audio input — text only.
  2. Health/Medicine-Enhanced: Optimized for medical knowledge reasoning and clinical safety. Outputs are informational and not a substitute for professional medical advice.
  3. MoE Architecture: 124B total parameters with 5.1B active per token — efficient inference despite large capacity.
  4. Reasoning Always On: Reasoning is always on for Ling 3.0 Flash Sante and cannot be disabled.
  5. Function Calling: Supports function calling via the tools and tool_choice inputs.
  6. Context & Output: 256K context window, 32K max output tokens.

Enjoy .....!

About
Ling 3.0 Flash Sante is a health and medicine-enhanced MoE model from InclusionAI with 124B total parameters and 5.1B activated per token. Optimized for medical knowledge reasoning, clinical safety, and evidence-based retrieval. Supports function calling with a 256K token context window.
About the Author
Novita AI is a serverless GPU cloud platform providing fast, cost-efficient API access to a wide catalog of open-source AI models. It offers OpenAI-compatible endpoints, flexible auto-scaling infrastructure, and pay-as-you-go pricing optimized for production AI workloads. Developers can integrate hundreds of models with a single API key without managing infrastructure.
Countries & Providers

Geographic locations of the external API endpoints this circuit connects to.

Novita AI
United StatesUS-CA
GermanyDE
IcelandIS
United KingdomGB
SingaporeSG
Pricing Structure

The cost model for running this circuit. Charges may be fixed per execution or based on token usage.

Token-based Pricing
UsageLocal
Input tokens0 credits
Output tokens0 credits
Token pricing display: Per TokenPer Token