- Cloudflare-hosted
- Function calling
- Reasoning
NVIDIA Nemotron 3 Super is a hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.
| Model Info | |
|---|---|
| Context Window ↗ | 256,000 tokens |
| Terms and License | link ↗ |
| Function calling ↗ | Yes |
| Reasoning | Yes |
| Unit Pricing | $0.50 per M input tokens, $1.50 per M output tokens |
Nemotron supports normal reasoning, low reasoning, and reasoning turned off. It does not currently select these modes from the top-levelreasoning_effort field. Pass the corresponding options through chat_template_kwargs instead.
To use low reasoning, set both enable_thinking andlow_effort:
response = client.chat.completions.create(
model="@cf/nvidia/nemotron-3-120b-a12b",
messages=[{"role": "user", "content": "What is the capital of Japan?"}],
max_tokens=16000,
temperature=1.0,
top_p=0.95,
extra_body={
"chat_template_kwargs": {
"enable_thinking": True,
"low_effort": True,
}
},
)The low-effort option appends {reasoning effort: low} to the latest user message. To turn reasoning off, setenable_thinking to False. For coding agents, setforce_nonempty_content to True in the samechat_template_kwargs object.
For more information, refer to NVIDIA's Nemotron API client example.
Try out this model with Workers AI LLM Playground. It does not require any setup or authentication and is an instant way to preview and test a model directly in the browser.
Launch the LLM Playground
export interface Env {
AI: Ai;
}
export default {
async fetch(request, env): Promise<Response> {
const messages = [
{ role: "system", content: "You are a friendly assistant" },
{
role: "user",
content: "What is the origin of the phrase Hello, World",
},
];
const stream = await env.AI.run("@cf/nvidia/nemotron-3-120b-a12b", {
messages,
stream: true,
});
return new Response(stream, {
headers: { "content-type": "text/event-stream" },
});
},
} satisfies ExportedHandler<Env>;
export interface Env {
AI: Ai;
}
export default {
async fetch(request, env): Promise<Response> {
const messages = [
{ role: "system", content: "You are a friendly assistant" },
{
role: "user",
content: "What is the origin of the phrase Hello, World",
},
];
const response = await env.AI.run("@cf/nvidia/nemotron-3-120b-a12b", { messages });
return Response.json(response);
},
} satisfies ExportedHandler<Env>;
import os
import requests
ACCOUNT_ID = "your-account-id"
AUTH_TOKEN = os.environ.get("CLOUDFLARE_AUTH_TOKEN")
prompt = "Tell me all about PEP-8"
response = requests.post(
f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/nvidia/nemotron-3-120b-a12b",
headers={"Authorization": f"Bearer {AUTH_TOKEN}"},
json={
"messages": [
{"role": "system", "content": "You are a friendly assistant"},
{"role": "user", "content": prompt}
]
}
)
result = response.json()
print(result)
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/nvidia/nemotron-3-120b-a12b \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{ "messages": [{ "role": "system", "content": "You are a friendly assistant" }, { "role": "user", "content": "Why is pizza so good" }]}'Input
stringID of the model to use (for example, '@cf/nvidia/nemotron-3-120b-a12b').objectParameters for audio output. Required when modalities includes 'audio'.number | nullPenalizes new tokens based on their existing frequency in the text so far.object | nullModify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.boolean | nullWhether to return log probabilities of the output tokens.integer | nullHow many top log probabilities to return at each token position (0-20). Requires logprobs=true.integer | nullThe maximum number of tokens to generate.integer | nullAn upper bound for the number of tokens that can be generated for a completion.object | nullSet of key-value pairs that can be attached to the object.array | nullOutput types requested from the model.integer | nullHow many chat completion choices to generate for each input message.booleandefault: trueWhether to enable parallel function calling during tool use.objectnumber | nullPenalizes new tokens based on whether they appear in the text so far.objectNemotron chat-template controls for normal reasoning, low-effort reasoning, and non-reasoning responses.one ofSpecifies the format the model must output.integer | nullIf specified, the system will make a best effort to sample deterministically.one ofboolean | nullWhether to store the output for model distillation or evaluation.boolean | nullIf true, partial message deltas will be sent as server-sent events.objectnumber | nullSampling temperature between 0 and 2.one ofControls which (if any) tool is called by the model. 'none' = no tools, 'auto' = model decides, 'required' = must call a tool.arrayA list of tools the model may call.number | nullNucleus sampling: considers the results of the tokens with top_p probability mass.stringA unique identifier representing your end-user, for abuse monitoring.objectOptions for the web search tool (when using built-in web search).one ofarrayminItems: 1maxItems: 128stringrequiredminLength: 1The input text prompt for the model to generate a response.Output
Synchronous — Send a request and receive a complete response
stringA unique identifier for the chat completion.stringintegerUnix timestamp (seconds) of when the completion was created.stringThe model used for the chat completion.arrayminItems: 1objectstring | nullStreaming — Send a request with `stream: true` and receive server-sent events
stringtext/event-streambinary