When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·28d ago
Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill and decode stages.
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs | NVIDIA Technical Blog
Nvidia Developer Blogby Jonathan Bentz·28d ago
Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software platform.
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·8 Sept 2026
In September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and NVIDIA will be growing and maturing CUDA...
Building a Memory-Driven Agent with NVIDIA NemoClaw | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·4 Sept 2026
Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it before contributing.
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·4 Sept 2026
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run locally on edge hardware.
How to Carry User Identity Across Federated Kubernetes and AI Platforms | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·3 Sept 2026
Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook where that data resides, and invoke...
NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·3 Sept 2026
AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents.
The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·2 Sept 2026
NVIDIA CUDA remains the foundation of GPU-accelerated computing, powering everything from scientific simulations to large-scale AI training.
Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·2 Sept 2026
This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and offers five guidelines for selecting...
Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·1 Sept 2026
AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons.
How to Size GPUs for AI Inference and TCO Without Overspending | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·1 Sept 2026
The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently size GPU resources for inference...
Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·31 Aug 2026
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next.
Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse NuRec | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·31 Aug 2026
A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle variant in the portfolio—and its...
Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·28 Aug 2026
Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing, post-processing, and runtime code.
NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure | NVIDIA Technical Blog
Nvidia Developer Blogby Farshad Ghodsian·26 Aug 2026
AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are...
How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents | NVIDIA Technical Blog
Nvidia Developer Blogby Tanya Lenz·26 Aug 2026
Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to continuously localize the robot...
Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·26 Aug 2026
Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·26 Aug 2026
Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.
Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo | NVIDIA Technical Blog
Nvidia Developer Blogby Michelle Horton·25 Aug 2026
When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels, and capturing NVIDIA CUDA graphs.
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access | NVIDIA Technical Blog
Nvidia Developer Blogby Elizabeth Goodman·25 Aug 2026
For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and maintain bindings back to Python...
Your filters hide everything on this page. Adjust them in preferences.