Trendora

NVIDIA Nemotron 3 Embed

Trial

Tools

An open and commercially available collection of embedding models for high-quality retrieval.

Why it's here

Placed in Trial: 8 article(s) of evidence from 4 source(s), led by model releases, with 5 in the last 30 days. Confidence 76%.

Evidence (8)

  • 8The New Stack·8/11/2026model_release
    NVIDIA launches Nemotron 3.5 Lightning and NeMo Switchyard

    NVIDIA expanded its Nemotron model family with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model optimized for high-volume agentic workloads. It also released NeMo Switchyard, an open-source routing library that directs requests to the most suitable model across open, proprietary, and NVIDIA systems without requiring app rewrites.

  • 6The New Stack·7/23/2026research
    Nvidia Says Local and Frontier Models Will Work Together

    Nvidia executive Joey Conway said organizations will increasingly use local/open models alongside frontier models, with routing systems choosing the best model for each task. He also pointed to enterprise control, lower latency, and lower cost as reasons to run adapted open models near the data, citing Nvidia hardware and software used to serve and orchestrate these workloads.

  • 7NVIDIA GenAI·7/17/2026research
    NVIDIA Vera Rubin Targets Better Intelligence per Dollar for Agentic AI Training

    NVIDIA argues that post-training has become the central workload in the agentic AI era, where models continuously adapt through reinforcement learning, tool use, and recovery from failures. The company highlights NeMo libraries and its Nemotron 3 Ultra open-weight model as examples of infrastructure and model design aimed at improving intelligence per dollar, with lower cost per token also improving inference economics.

  • 7Hacker News·7/16/2026model_release
    German consortium releases Soofi S, an open 30B model

    A German research consortium led by the KI Bundesverband has released Soofi S 30B-A3B, an open language model trained on Deutsche Telekom's Industrial AI Cloud. The model uses a hybrid Mamba-Transformer mixture-of-experts design and is reported to outperform other fully open models on English, German, and programming benchmarks, while maintaining high throughput on long contexts.

  • 8Hugging Face Blog·7/16/2026model_release
    NVIDIA Nemotron 3 Embed Tops RTEB

    NVIDIA released Nemotron 3 Embed, an open embedding model collection aimed at improving retrieval for enterprise RAG, agentic workflows, code retrieval, and memory. The 8B BF16 model ranks #1 on the RTEB leaderboard, while 1B variants target lower-cost and higher-throughput production deployment.

  • 8NVIDIA GenAI·5/21/2026model_release
    NVIDIA Debuts Nemotron 3 Ultra for Long-Running AI Agents

    At NVIDIA GTC Taipei at COMPUTEX, NVIDIA announced Nemotron 3 Ultra, an open 550-billion-parameter mixture-of-experts model designed for long-running AI agents. The company says it delivers up to 5x faster inference and can reduce the cost of complex agentic tasks by up to 30%, with early adopters including Perplexity, Palantir and ServiceNow.

  • 8Hugging Face Blog·4/28/2026model_release
    NVIDIA Launches Nemotron 3 Nano Omni for Long-Context Multimodal AI

    NVIDIA introduced Nemotron 3 Nano Omni, an omni-modal model for document analysis, speech recognition, long audio-video understanding, and agentic computer use. The model claims strong benchmark results on document, video, audio, and voice tasks, along with higher throughput and efficiency than comparable open multimodal models.

  • 7Hugging Face Blog·1/5/2026product_launch
    NVIDIA shows DGX Spark and Reachy Mini AI agent demo

    NVIDIA used CES 2026 to showcase a demo that runs an agent locally on DGX Spark and interacts through the Reachy Mini robot. The Hugging Face blog post explains how to reproduce the setup with open NVIDIA models, an agent toolkit, and optional local, cloud, or serverless deployment paths.