Introducing NVIDIA Nemotron 3.5 Lightning⚡
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster.
It delivers up to 4x the output speed of similar-sized models.
Another leaderboard win for Nemotron 🏆
Nemotron 3 Embed 8B ranks #1 for combined nDCG@10 on the Q2D-Web benchmark, tested across 190M web documents and nearly 70K agent-reformulated queries in 10 languages.
Shoutout to @perplexity_ai for putting it together.
We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems.
Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries.
Read more: perplexity.ai/hub/blog/q2d-w…
Most diffusion-based video models are trained on short clips. When they’re used to relight longer videos in chunks, the lighting can visibly jump between them.
Meet HorizonRelight, new work from our research team and @USC. It propagates target-domain context from one sliding
Multimodal models put different demands on vision encoding, prefill and decoding.
Separating vision encoding from the other stages can reduce resource contention and help models respond faster, but only for the right workloads.
See how EPD disaggregation works, when it helps