Skip to content
Dashboard

Nemotron 3.5 Lightning 30B

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.

Input and output price
Prices from: Input $0.05, Output $0.15, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'nvidia/nemotron-3.5-lightning',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by NVIDIA

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M0.3 s211 tps
$0.50/M
$2.40/M
Read$0.12/M
baseten logo
deepinfra logo
fireworks logo
+1
06/04/2026
256K0.2 s172 tps
$0.15/M
$0.65/M
bedrock logo
03/11/2026
262K0.2 s135 tps
$0.05/M
$0.24/M
deepinfra logo
12/15/2025
131K0.2 s94 tps
$0.20/M
$0.60/M
bedrock logo
deepinfra logo
10/28/2025
131K0.2 s147 tps
$0.06/M
$0.23/M
bedrock logo
deepinfra logo
08/18/2025

Your use is subject to NVIDIA's Terms & Privacy Policies.