The model catalog.
Every model in the catalog, with its tier, context window, capabilities, and per-million-token list pricing. Request one by name with the model field. Every paid plan includes the whole catalog; the Free plan includes the open-model lineup. This is a snapshot; GET /v1/models is always the live source of truth.
sdn-opus-5The frontier ceiling — deepest reasoning and vision for the hardest problems.
sdn-opus-4.7Prior Opus flagship — frontier reasoning, pinned in place.
sdn-gpt-5.5Frontier generalist — broad, sharp, and steady on long tool chains.
sdn-gpt-5.4The prior GPT-5 flagship — most of 5.5's strength for less.
sdn-gemini-3-proFrontier scale — a 1M-token window and strong multimodal reasoning.
sdn-kimi-k3Moonshot's flagship — deep reasoning, vision, and a 1M-token window.
sdn-deepseek-v4-proDeepSeek's flagship — frontier-grade reasoning at a startling price.
sdn-qwen3.7-maxThe current Qwen flagship — frontier scale with a 1M-token window.
sdn-inklingThinking Machines' frontier debut — text, images, and audio in one model.
sdn-gpt-oss-120bDefault. Strong general-purpose agent model with fast responses.
sdn-sonnet-4.6The premium ceiling — top reasoning and vision for the hardest work.
sdn-sonnet-5The newest Sonnet — near-frontier quality for everyday premium work.
sdn-haiku-4.5The fast Claude — premium-family quality at a fraction of the latency.
sdn-grok-4.3Deep reasoning over a very large window, at the low end of premium.
sdn-glm-5.2Highest-quality GLM. Flagship tuned for quality over raw speed — higher, variable latency.
sdn-glm-5Fast, reliable coding flagship for everyday heavy work.
sdn-qwen3-coder-480bHeavy coding flagship — large MoE built for complex code.
sdn-kimi-k2.7-codeCoding-specialist flagship. Quality-first; higher, variable latency.
sdn-qwen3-maxThe prior Qwen flagship — a heavyweight that codes exceptionally well.
sdn-qwen3-235bBig reasoning model at a cheap-tier price — standout value.
sdn-deepseek-v3.2Strong reasoning. Best on open-ended analysis, not strict tool loops.
sdn-kimi-k2-thinkingExtended reasoning; budget output tokens for its hidden chain-of-thought.
sdn-minimax-m2.5Cheap reasoning with a large context window.
sdn-ernie-x1Baidu's reasoning specialist — deep deliberation at a low price.
sdn-hunyuan-t1Tencent's reasoning model — strong long-form thinking, very cheap.
sdn-step-3StepFun's multimodal reasoner — thinks, sees, and calls tools.
sdn-qwen3.5-397bQwen's biggest open-weight reasoner — flagship thinking, mid-tier price.
sdn-nemotron-3-ultraThe top Nemotron — 550B of deliberate reasoning at a mid-tier price.
sdn-nemotron-super-3-120bStrong all-round workhorse with a very large context.
sdn-nemotron-3-120bHybrid MoE, strong on multi-agent, 256K context.
sdn-qwen3-next-80bEfficient workhorse — big context at a low price.
sdn-gemma-4-31bGoogle's open workhorse — reasoning and a 256K window near the price floor.
sdn-mistral-large-3-675bLarge general-purpose model, fast and capable.
sdn-ernie-5.1Baidu's flagship generalist — broad knowledge with vision.
sdn-hunyuan-turbosTencent's fast generalist — quick answers at a rock-bottom price.
sdn-gemini-3-flashGoogle's fast frontier model — 1M context, multimodal, cheap.
sdn-minimax-m3MiniMax's frontier agent model — 1M context and vision at a workhorse price.
sdn-kimi-k2.6Kimi's vision generalist — thinks when asked, sees what you show it.
sdn-glm-4.7The GLM workhorse — flagship instincts at an everyday price.
sdn-qwen3.7-plusQwen's balanced mid-tier — a 1M-token window at an everyday price.
sdn-qwen3-vl-235bAffordable eyes — a 235B vision model at workhorse money.
sdn-hunyuan-hy3Tencent's newest generalist — reasoning and tools near the price floor.
sdn-llama-3.3-70bMeta's dependable open workhorse — a known quantity everywhere.
sdn-qwen3-coder-30bCheap coding offload for routine changes.
sdn-qwen3-coder-nextBalanced coding model with a large context.
sdn-kimi-k2.5Fast coding-general model.
sdn-devstral-2-123bCoding-specialist tuned for software tasks.
sdn-minimax-m2.7MiniMax's coding workhorse — agentic edits at a budget rate.
sdn-kat-coder-proKuaishou's coding specialist — direct edits, no reasoning overhead.
sdn-gemma-4-26b256K context for the price of a small model.
sdn-longcat-2Meituan's flagship MoE — a 1M-token window with an unusually deep cached-input discount.
sdn-deepseek-v4-flashThe V4 line's fast sibling — a 1M-token window at cheap-tier prices.
sdn-qwen3.5-flashA million tokens of context at one flat, tiny price.
sdn-mimo-v2.5Xiaomi's efficiency play — 1M context and vision for pocket change.
sdn-qwen-turboThe catalog's cheapest chat tokens — quick, direct answers at volume.
sdn-qwen3.8-maxFlagship reasoning with images and a near-million-token window.
sdn-gpt-oss-20bSmallest and fastest tier.
sdn-glm-4.7-flashNear the price floor — GLM-family quality in the budget tier.
sdn-nemotron-nano-3-30bCheapest tool + reasoning capable model.
sdn-step-3.7-flashStepFun's quick multimodal — sees, thinks, and answers fast for very little.
sdn-bge-m3BAAI's versatile embedding — multilingual, multi-granularity.
sdn-bge-base-enThe long-standing default English embedding.
sdn-bge-small-en384 dimensions — the smallest index and the fastest search.
sdn-bge-large-enThe most accurate English embedding in the catalog.
sdn-embed-gemma-300mCompact Gemma-family embedding with a 2K input window.
sdn-qwen3-embed-0.6bMultilingual retrieval with an 8K input window.
sdn-plamo-embed-1bJapanese-specialist embedding — the widest vector in the catalog.
Prices are USD per one million tokens — the public list rates your plan's included usage is measured at. Click any card for the model's full page: description, strengths, and capabilities.
Call GET /v1/models to read the catalog programmatically — each entry is self-describing, including capabilities and pricing. The gateway is always authoritative over this snapshot.