OpenDLSS Brings NVIDIA DLSS 5 to Vulkan with Bit-Exact Precision
AI News

OpenDLSS Brings NVIDIA DLSS 5 to Vulkan with Bit-Exact Precision

6 min
10/1/2026
OpenDLSSVulkanDLSS 5Neural Rendering

OpenDLSS Brings NVIDIA's DLSS 5 to Vulkan with Bit-Exact Precision

In a significant move for the graphics community, a developer known as maanHimself has released OpenDLSS-NR, a Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network. The project, now available on GitHub, claims to be bit-exact against the original, meaning it produces identical output to NVIDIA's proprietary implementation—down to the last byte. This is a major achievement that could democratize access to AI-driven neural rendering across GPU vendors.

The repository, which has already garnered 367 stars and 38 forks, provides a complete implementation of the DLSS 5 neural network—a 71-block Swin/ViT architecture with six pooling levels. It runs on Vulkan, making it compatible with any GPU that supports the required tensor operations, including AMD Radeon and Intel Arc, not just NVIDIA's RTX lineup.

What Is DLSS 5 Neural Rendering?

DLSS 5 represents a shift from traditional upscaling to generative neural rendering. Unlike earlier DLSS versions that primarily focused on resolution scaling, DLSS 5 re-renders the frame entirely. It takes a low dynamic range proxy of the rendered frame, injects Gaussian noise, and generates new detail—adjusting tone, structure, and even skin textures based on a style setting. The network outputs an RGB residual and a temporal-blend logit at the same resolution as the input; it is not an upscaler but a full re-rendering network.

This generative approach is what makes the reimplementation so challenging. The network operates in FP8 (E4M3) with FP16 accumulation, requiring 141 MiB of weights. It processes 241 dispatches per frame, making performance a critical factor for real-time use.

Technical Milestones and Design

OpenDLSS-NR achieves its bit-exactness through a combination of techniques. The project includes both GLSL reference kernels and optimized PTX kernels for NVIDIA GPUs. The PTX route leverages mma.sync instructions for FP8 matrix multiplication, cp.async rings for efficient memory transfer, and barrier-free chaining through device counters to minimize synchronization overhead.

For non-NVIDIA hardware, the GLSL route provides a fallback that materializes every intermediate, ensuring correctness at the cost of performance. The project also includes a WebGPU port that runs the same network in a browser without tensor cores or FP8 support, demonstrating the portability of the implementation. The WebGPU version runs at 72 ms at 512x512, compared to 2.7 ms on native Vulkan with RTX hardware.

The developer emphasizes that the exactness lies in the specification, not the hardware. By documenting the network graph, weight layouts, and numerical contracts, OpenDLSS-NR enables other implementations to achieve the same results.

Performance Benchmarks

On an RTX 4070 SUPER, OpenDLSS-NR delivers impressive performance. The entire network runs in:

  • 2.8 ms at 768x768
  • 7.8 ms at 1920x1080
  • 12.6 ms at 2560x1440
  • 29.3 ms at 3840x2160

These figures represent the minimum over 40 frames, with medians running slightly higher due to GPU clock state changes. The performance is competitive with NVIDIA's proprietary implementation, making OpenDLSS-NR a viable drop-in replacement for developers.

continue reading below...

How It Works: The Network Architecture

The DLSS 5 network is a U-net of shifted-window transformer blocks with a global vision transformer (ViT) at the bottleneck. It takes multiple inputs: a rendered frame (as a low dynamic range proxy), three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars. The output is four f32 channels per pixel—an RGB residual and a temporal blend logit.

The project includes comprehensive documentation in its docs/ folder, covering the network graph, numerical exactness contract, weight layouts, execution scheduling, and the demo's frame loop. NVIDIA's own research report on DLSS 5 is referenced as the authoritative source for the model architecture.

Legal and Licensing Considerations

OpenDLSS-NR is released under the MIT license, making it freely available for commercial and personal use. The project contains no NVIDIA software, weights, headers, or instructions for obtaining them. Users must supply their own model weights, which the developer notes are not included in the repository.

The project explicitly states it is not affiliated with or endorsed by NVIDIA. It grants no rights under any NVIDIA intellectual property, and users are responsible for ensuring they have the right to use whatever model data they load. This approach sidesteps potential legal issues while still providing a functional implementation.

Why This Matters

The release of OpenDLSS-NR could have significant implications for the graphics industry. By providing an open-source, Vulkan-compatible implementation of DLSS 5, it enables:

  • Cross-platform support: Games using OpenDLSS-NR could run on AMD, Intel, and NVIDIA GPUs, as well as on Linux and potentially consoles that support Vulkan.
  • Modding and research: Developers and researchers can study and modify the network without being locked into NVIDIA's proprietary stack.
  • Competition: AMD and Intel could potentially adopt OpenDLSS-NR as a foundation for their own neural rendering solutions, reducing NVIDIA's competitive advantage.

The project is still in its early stages—it currently supports Windows with NVIDIA Ada or newer GPUs for the fastest path, but the WebGPU port shows potential for broader reach. As the community contributes, we may see support for more platforms and hardware.

Getting Started

For developers interested in trying OpenDLSS-NR, the build process is well-documented. The project uses a portable toolchain that downloads glslang, Vulkan-Headers, volk, CMake, and Ninja—no Vulkan SDK installation required. A demo application using Filament is included, along with tools for benchmarking, profiling, and verifying bit-exactness against recorded fixtures.

The developer has also provided a parity command that compares output against captures of the original DLSS 5, ensuring that the implementation remains faithful. This rigorous approach to validation is a hallmark of the project's quality.

Conclusion

OpenDLSS-NR is a technical tour de force that demonstrates the power of reverse engineering and open-source collaboration. By reimplementing NVIDIA's DLSS 5 neural rendering network in Vulkan with bit-exact precision, it opens up new possibilities for cross-platform AI graphics. While it may not immediately replace NVIDIA's proprietary solution, it provides a solid foundation for innovation and competition.

As the project evolves, we can expect to see broader hardware support, performance improvements, and potentially integration into game engines. For now, it stands as a testament to what's possible when developers take on the challenge of unlocking proprietary technology for the benefit of all.