A Windows desktop app (C++20, Win32 + Direct3D 11, Dear ImGui) that upscales
images with classic resampling or neural models loaded directly from .pth
files, running on the GPU via DirectML (any vendor) — no Python, no
offline conversion.
Open Source/Clarity/Clarity.slnx in Visual Studio and build x64.
A .pth holds only weights, not a network graph. Clarity:
- reads the PyTorch file in C++ (PthLoader.cpp) — unzips the container and interprets the pickle to recover the weight tensors;
- auto-detects the architecture from the tensor names/shapes and builds an ONNX graph in C++ from the weights (NeuralEngine.cpp + a tiny hand-written ONNX serializer, OnnxBuilder.cpp);
- runs it through ONNX Runtime with the DirectML execution provider (GPU on NVIDIA/AMD/Intel; CPU fallback), tiled.
No Python and no protobuf/onnx libraries — the only external SDK is ONNX Runtime.
Download the ONNX Runtime Windows x64 DirectML build:
- NuGet
Microsoft.ML.OnnxRuntime.DirectML, or - GitHub release
onnxruntime-win-x64-directml-<version>.zip.
Lay it out as:
Source/SDKs/onnxruntime/include/ (onnxruntime_cxx_api.h, dml_provider_factory.h, …)
Source/SDKs/onnxruntime/lib/ (onnxruntime.lib, onnxruntime.dll, DirectML.dll)
The VS project auto-detects this (x64), defines CLARITY_ENABLE_ONNX, links
onnxruntime.lib, and copies the DLLs next to the exe. Works in Debug and
Release x64. Without it, the app still builds with classic upscaling only.
- Open / drag-drop an image; classic upscaling (1×–8×) works with no SDK.
- Method → Neural (.pth AI), then Import .pth… (or drag a
.pthon the window; any.pthinData/is auto-listed). The app detects the model, shows the architecture, and whether it's running on GPU / DirectML or CPU.
- ESRGAN / RRDBNet family, ×4 — modern Real-ESRGAN naming and classic
"old-arch" ESRGAN (
RealESRGAN_x4plus,RealESRNet_x4plus, ESRGAN model-database.pth, …). Runs with dynamic tile sizes. - SwinIR (×2 / ×4,
nearest+convupsampler) — including the two bundledData/*.pth(SwinIR-M and SwinIR-L real-world SR). The window-attention transformer is emitted as a static ONNX graph built for a fixed tile size; relative-position bias and shifted-window masks are precomputed on the CPU and baked in as constants.
Not yet: RRDBNet ×2 / ×1 (input pixel-unshuffle), SwinIR pixelshuffle
upsampler (classical-SR variants).
- Its tile size is fixed when the model loads (a transformer needs a static shape). To change it, move the Tile size slider then re-select the model. Smaller tiles = less memory and faster per tile; SwinIR-L is heavy, so start small (e.g. 128) and on a modest image.
- Some 5-D/6-D reshape/transpose nodes may fall back to CPU inside DirectML; that's correctness-preserving but can slow things down.
Data/ drop .pth models here (auto-listed)
Source/Clarity/Clarity.slnx Visual Studio solution (build x64)
Source/Clarity/Clarity/
main.cpp Win32 + D3D11 bootstrap, window, drag-drop
App.{h,cpp} UI + application logic, threaded jobs
Image.{h,cpp} RGBA8 image + stb load/save
Upscaler.{h,cpp} classic resampling (stb_image_resize2)
Tensor.h float weight container
PthLoader.{h,cpp}native .pth (zip + pickle) reader
OnnxBuilder.{h,cpp} in-memory ONNX model writer (hand-rolled protobuf)
NeuralEngine.{h,cpp} arch detection, ONNX graph build, ORT + DirectML, tiling
stb_impl.cpp single TU compiling the stb libraries
Source/SDKs/onnxruntime you add this (DirectML build) to enable neural