## Short Segments Today, NVIDIA's Nemotron 3 Embed model collection is redefining AI retrieval capabilities. The newly released models, including the standout 8B checkpoint, are designed to enhance retrieval accuracy for AI systems across various applications. Coming up, we'll explore how these models are setting new benchmarks and what this means for developers and enterprises. ## Feature Story NVIDIA's Nemotron 3 Embed model collection is making waves in the AI community by setting a new standard in retrieval accuracy. The collection, which includes three open checkpoints, is designed to improve the retrieval capabilities of AI systems, particularly in production-scale retrieval-augmented generation (RAG), agentic retrieval, code retrieval, and agent memory. The flagship model, Nemotron-3-Embed-8B-BF16, has achieved the top rank on the Retrieval Embedding Benchmark (RTEB), a significant achievement that highlights its superior performance. This model, along with its smaller counterparts, Nemotron-3-Embed-1B-BF16 and Nemotron-3-Embed-1B-NVFP4, offers developers a range of options to balance accuracy and efficiency in their AI applications. All models in the collection are transformer encoders trained with bidirectional attention masking, and they utilize average pooling over token-level representations to generate the final embeddings. With a maximum sequence length of 32,768 tokens, these models are equipped to handle extensive data inputs, making them ideal for complex retrieval tasks. One of the standout features of the Nemotron 3 Embed models is their multilingual capability. Evaluated across 34 languages, these models are built on Mistral bases, with the 8B model using the Ministral-3-8B-Instruct-2512 and the 1B variants using the Ministral-3-3B-Instruct-2512. This multilingual support ensures that the models can be effectively deployed in diverse linguistic contexts, broadening their applicability. The release of these models is particularly significant for developers and enterprises looking to enhance their AI systems' retrieval accuracy. By providing open and commercially available models, NVIDIA is enabling developers to access state-of-the-art retrieval technology without the constraints of licensing fees. This openness not only democratizes access to advanced AI capabilities but also fosters innovation by allowing developers to build upon and customize the models for their specific needs. In practical terms, the Nemotron 3 Embed models are poised to improve the performance of AI systems in various domains. For instance, in enterprise search, these models can enhance the accuracy and relevance of search results, leading to more efficient information retrieval. In code retrieval, they can assist developers in finding relevant code snippets more quickly, streamlining the development process. Additionally, in agent memory applications, the models can help AI systems maintain and retrieve long-term information more effectively. The availability of these models on platforms like Baseten further underscores their accessibility and ease of deployment. Developers can now integrate the Nemotron 3 Embed models into their existing workflows, leveraging their high retrieval accuracy to improve the overall performance of their AI systems. Looking ahead, the impact of the Nemotron 3 Embed models is likely to be felt across the AI landscape. As more developers adopt these models, we can expect to see improvements in the accuracy and reliability of AI systems, particularly in tasks that require precise information retrieval. This development not only enhances the capabilities of AI agents but also sets a new benchmark for retrieval performance in the industry. In conclusion, NVIDIA's release of the Nemotron 3 Embed model collection marks a significant advancement in AI retrieval technology. By offering open, high-performance models that excel in multilingual contexts, NVIDIA is empowering developers to build more accurate and efficient AI systems. As these models gain traction, they are set to transform the way AI systems retrieve and process information, paving the way for more intelligent and capable AI applications.