All models
Model card

Nemotron 3 Ultra

sdn-nemotron-3-ultraReasoningLive in the catalog

NVIDIA's largest Nemotron: a 550B mixture-of-experts (55B active) that reasons on demand and keeps its tool output clean, priced well under what its scale suggests. The natural step up when the Super 120B runs out of depth but the task doesn't justify frontier rates.

Context window202K tokens
Input · per Mtok$0.6
Output · per Mtok$2.4
Served bySideren engine
Where it earns its keep[01/03]
  • 01550B-scale reasoning on demand
  • 02Clean structured tool output
  • 03202K context at a mid-tier price
Capabilities
Tool callingYes

Strict, schema-faithful tool calls, enforced by the engine on every request — safe to build an agent loop on.

ReasoningYes

Thinks before it answers. Budget max_tokens generously — hidden reasoning counts against it.

VisionNo

Text-only.

Behind the endpoint[02/03]

One endpoint. Served by our engine.

Nemotron 3 Ultra is served through the Sideren engine — zero-downtime serving is the design target, not a status-page apology. You request it by name; everything else is our problem.

Call it by name
curl https://api.sideren.io/v1/messages \
  -H "x-api-key: sdn_your_key" \
  -H "content-type: application/json" \
  -d '{
    "model": "sdn-nemotron-3-ultra",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

OpenAI-style clients work too — send the same model name to /v1/chat/completions with a Bearer key. See the docs for both dialects.

Put your agent on real infrastructure

Your agent doesn't change.
Everything underneath does.

$export ANTHROPIC_BASE_URL=https://api.sideren.io

Start free · no card · Starter from $5/mo