NVIDIA Nemotron Nano 12B v2 VL BF16 API
surplus/nvidia-nemotron-nano-12b-v2
Call NVIDIA Nemotron Nano 12B v2 VL BF16 through UberLLM’s OpenAI-compatible API — from $0.2000/M input and $0.6000/M output, routed automatically to the cheapest healthy provider. NVIDIA · 128,000 token context. No inference markup.
NVIDIA Nemotron Nano 12B v2 VL BF16 pricing & providers
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| Surplus Intelligencebest | $0.2000 | $0.6000 |
How to use NVIDIA Nemotron Nano 12B v2 VL BF16 with the OpenAI SDK
Already using OpenAI or OpenRouter? Change one line — the base URL — and call NVIDIA Nemotron Nano 12B v2 VL BF16 with your existing code. Anthropic SDKs work against https://api.uberllm.dev/v1/messages.
from openai import OpenAI
client = OpenAI(
base_url="https://api.uberllm.dev/v1",
api_key="ull_your_key_here",
)
resp = client.chat.completions.create(
model="surplus/nvidia-nemotron-nano-12b-v2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)NVIDIA Nemotron Nano 12B v2 VL BF16 API — FAQ
How much does the NVIDIA Nemotron Nano 12B v2 VL BF16 API cost?
NVIDIA Nemotron Nano 12B v2 VL BF16 costs from $0.2000 per million input tokens and $0.6000 per million output tokens through UberLLM, which routes to the cheapest healthy provider. There is no inference markup.
How do I call NVIDIA Nemotron Nano 12B v2 VL BF16 via API?
Use any OpenAI-compatible SDK: set the base URL to https://api.uberllm.dev/v1, your UberLLM API key, and model "surplus/nvidia-nemotron-nano-12b-v2". Anthropic SDKs work via https://api.uberllm.dev/v1/messages.
Is NVIDIA Nemotron Nano 12B v2 VL BF16 OpenAI-compatible on UberLLM?
Yes. Every model on UberLLM — including NVIDIA Nemotron Nano 12B v2 VL BF16 — is served through one OpenAI-compatible endpoint, so you change only the base URL and key.