Ling 3.0 Flash Fin is a finance-enhanced MoE model from InclusionAI with 124B total parameters and 5.1B activated per token. Optimized for investment workflows, multi-step reasoning, and long-horizon planning. Supports function calling with a 256K token context window — served via Novita AI serverless inference. This circuit provides an OpenAI-compatible chat completion endpoint backed by the inclusionai/ling-3.0-flash-fin model.
messages)arraymax_tokens)number32768temperature)number1top_p)number0.95tools)array[]tool_choice)auto or none.stringautochat_completion_output)choices[0].delta.content stream chunks.reasoning_output)choices[0].delta.reasoning_content stream chunks.chat_completion_tool_calls)choices[0].delta.tool_calls stream chunks.You can try this circuit directly in the Playground — no setup required.
tools and tool_choice inputs.Geographic locations of the external API endpoints this circuit connects to.
The cost model for running this circuit. Charges may be fixed per execution or based on token usage.