Releases: nihui/ncnn
Release list
Release 20260526
Release 20260525
新增 layer 与核心特性
新增 ARM SDPA layer 实现,支持 ARM 平台 scaled-dot-product-attention 加速 (@Abandon-ht, Tencent#6698)
新增 Vulkan SDPA layer 与 unroll 2x2 / local memory 优化 (Tencent#6514)
SDPA Vulkan 增加 cooperative matrix flash attention 优化 (Tencent#6528)
SDPA Vulkan 增加无 cooperative matrix 的 flash attention 优化路径 (Tencent#6538)
SDPA Vulkan 灵活支持 fp16/bf16 coopmat,统一 qk 与 qkv 的 cross shader (Tencent#6521)
SDPA Vulkan 在 bf16s 路径下移除 cm,并新增 sdpa decode/prefill 性能测试 (@futz12, Tencent#6632)
新增 Vulkan rotaryembed layer (@futz12, Tencent#6519)
新增 Vulkan reduction shader (@futz12, Tencent#6476)
新增 Vulkan softplus shader (@futz12, Tencent#6478)
新增 Vulkan shrink shader (@futz12, Tencent#6479)
新增 Vulkan groupnorm 实现,含 reduce_mean、coeffs、norm 多阶段 shader (Tencent#6556)
新增 Vulkan unfold (im2col) 算子,含多 elempack shader 变体 (@futz12, Tencent#6543)
新增 ncnn2table 集成 Embed/MHA/RNN/LSTM/GRU 权重 scale 生成,简化 int8 量化流程 (@Roundaboutt, Tencent#6688)
新增 benchncnn_llm 大模型 benchmark 工具,内置 hunyuan-0.5b / llama3.2-1b / minicpm4-0.5b / qwen2.5-0.5b / qwen3-0.6b / tinyllama-1.1b / youtu-llm-2b 等 7 个 LLM 模型 param (Tencent#6711)
新增 Vulkan pipeline cache 单文件持久化保存与加载支持,含 device/driver/shader 严格校验、SPIR-V 持久化与文档 (@futz12, @CLV-Iclucia, Tencent#6702)
新增 use_mapped_model_loading 选项与 readonly mappedfile 包装类,支持 mmap 模型加载,显著降低 LLM 加载时内存占用 (Tencent#6537)
新增 use_weights_in_host_memory 选项与 vkhostallocator,发现并使用 VK_EXT_external_memory_host 扩展 (Tencent#6531)
新增 GPU resizable BAR 特性发现,启用后大幅提升权重上传速度 (Tencent#6536)
新增 Vulkan 按 layer 分批上传权重,降低加载阶段峰值显存占用 (Tencent#6534)
新增 vkmat 增加 memory type index 与 is_device_local 字段 (Tencent#6581)
新增 Vulkan 在收集到足够 dispatch 后及时提交 compute command,避免驱动 timeout (Tencent#6541)
扩展 unaryop 支持 sign / sinh / cosh / asinh / acosh / atanh / expm1 / log1p,覆盖 scalar、SIMD 与 Vulkan 路径,并打通 onnx2ncnn 与 pnnx 表达式转换 (@crafcat7, Tencent#6675)
扩展 binaryop 支持 fmod / logaddexp / floor_divide / remainder (@futz12, Tencent#6549)
通用层 absval / dropout / glu / cumulativesum / instancenorm / layernorm / rmsnorm / log / power / mvn / normalize / prelu / scale / shufflechannel 等扩展 4d Mat 输入输出支持,并补充测试与文档 (Tencent#6737)
benchncnn vision_transformer 改用 Gemm 实现线性层 (Tencent#6709)
benchncnn_llm 默认 seqlen 由 1024 改为 256,更接近真实推理场景 (Tencent#6738)
benchmark 模型文件迁移至 models 子目录 (Tencent#6710)
新增 perf 测试基础设施 tests/perf,覆盖 batchnorm / binaryop / concat / convolution / convolutiondepthwise / innerproduct / pooling / relu / sigmoid / softmax 等典型层 (Tencent#6570)
Vulkan 优化与重构
Vulkan convolution packed 统一 elempack,去除大量重复 shader (Tencent#6561)
Vulkan convolution 1x1s1d1 统一 elempack (Tencent#6562)
Vulkan convolution gemm 统一 elempack (Tencent#6565)
Vulkan deconvolution packed 统一 elempack (Tencent#6564)
Vulkan deconvolution gemm 统一 elempack (Tencent#6572)
Vulkan convolution1d packed 统一 elempack (Tencent#6566)
Vulkan conv1d 增加 cooperative matrix gemm 路径 (@futz12, Tencent#6587)
Vulkan gemm packed 实现 (Tencent#6573)
Vulkan gemm bf16 cooperative matrix 优化(输出与 a/b 不能共享 shared memory,bf16p 在 fp16s 上的优化)(Tencent#6515)
Vulkan gemm/sdpa unroll 4x4,向量化 a 与 b 的加载,避免 bank conflict,增加 subgroup 优化 (Tencent#6524)
Vulkan layer 增加 packed shape hint,简化 binaryop / cast / concat 等 30+ 层 (Tencent#6553)
Vulkan 加载模型时遵守 extension 依赖关系 (Tencent#6705)
Vulkan 模型加载时清理无效的 bf16 选项 (Tencent#6522)
Vulkan weight allocator 优先采用 host memory 融合分配 (Tencent#6545)
Windows WDDM 下 host 分配的内存导入后仍占共享显存,因此直接使用共享显存路径 (Tencent#6547)
Vulkan 修复
修复 binaryop vulkan shader 在 MoltenVK 上的编译错误 (Tencent#6602)
修复 Vulkan reduction 中 subgroup float16 扩展应作为必需依赖 (@NKID00, Tencent#6615)
绕过 SwiftShader 在 memory_type_bits 字段中错误设置 memory property flags 的问题 (Tencent#6539)
针对 Adreno GPU 驱动硬件限制,禁用 cooperative matrix (Tencent#6719)
绕过 LLVMpipe 在 atan2 零值时返回错误结果 (Tencent#6729)
x86 优化
x86 gemm 支持 bf16 storage (Tencent#6598)
x86 gemm bf16s 增加 avx512bf16 优化 (Tencent#6609)
x86 gemm int8 用 alignr 重排,提升性能 (Tencent#6600)
针对 AMD Zen 5 上 vpalignr 与 vdpbf16ps 端口冲突,将 BF16 GEMM micro kernel 中的 _mm*_alignr_epi8(x,x,N) 替换为 _mm*_shuffle_epi32(vpshufd),并对 16x16 kernel 增加指令调度,约 16% 性能提升 (Tencent#6673)
x86 conv bf16s avx512 unroll 16 (Tencent#6680)
x86/arm gemm 在 m==1 场景下专门优化 (Tencent#6723)
x86 fp16s innerproduct gemm 重排消除 loop-carried 依赖延迟 (@Edwardssss, Tencent#6682)
x86 int8 sse4.1 路径优化 innerproduct 与 convolutiondepthwise (@Edwardssss, Tencent#6687)
x86 PixelShuffle 用 SIMD block transpose 优化 (@crafcat7, Tencent#6690)
x86 erf 与 gelu 优化 (@futz12, Tencent#6604)
x86 interp 优化 (Tencent#6597)
x86 rotaryembed SIMD 优化 (@futz12, Tencent#6427)
x86 bf16/fp16 storage 大规模铺开
AbsVal_x86 增加 fp16s/bf16s 支持 (Tencent#6584)
LayerNorm_x86 bf16s 支持,含 avx512bf16 dispatch (Tencent#6585)
RMSNorm_x86 bf16s 支持,含 avx512bf16 dispatch (Tencent#6586)
UnaryOp_x86 bf16s 支持 (Tencent#6588)
clip / relu / sigmoid x86 bf16s 支持 (Tencent#6589)
BinaryOp_x86 bf16s 支持,含 avx512bf16 dispatch (Tencent#6591)
concat / slice / flatten / reshape / crop / padding / packing x86 fp16/bf16 storage 支持 (Tencent#6593)
groupnorm / instancenorm bf16s 支持,含 avx512bf16 dispatch (Tencent#6594)
batchnorm / prelu / scale / swish / softmax bf16s 支持,含 avx512bf16 dispatch (Tencent#6595)
gemm 增加 out_elemtype,multiheadattention 与 sdpa bf16s 支持,跳过 mha bf16 测试 (Tencent#6623)
rotaryembed / tanh / selu / mish / hardswish / hardsigmoid / gelu / erf / elu / eltwise / dropout / quantize / dequantize / bnll x86 bf16s 支持 (Tencent#6624)
innerproduct x86 bf16s 支持 (Tencent#6625)
convolution x86 bf16s 支持 (Tencent#6626)
deconvolution x86 bf16s 支持,整理头文件 (Tencent#6627)
convolution1d x86 bf16s 支持 (Tencent#6630)
pooling x86 bf16s 支持 (Tencent#6648)
interp x86 bf16s 支持 (Tencent#6649)
x86 重构
x86 deformableconv2d packed 统一 elempack,删除大量重复 pack16 头文件 (Tencent#6567)
x86 deconvolution packed 统一 elempack (Tencent#6568)
ARM 优化
armv8.4 bf16 gemm 优化 (Tencent#6714)
armv8.4 bf16 convolution im2col-gemm 优化 (Tencent#6715)
armv8.4 bf16 innerproduct 优化 (Tencent#6716)
arm multiheadattention 支持 bf16 storage (Tencent#6717)
arm 上 erf / elu / gelu / selu SIMD 加速 (@futz12, Tencent#6605)
aarch64 上 exp_ps 的 floor 步骤改用 vrndmq_f32,加速热点路径 (@crafcat7, Tencent#6657)
aarch64 上 fp16 exp_ps 的 floor 步骤同步加速 (@crafcat7, Tencent#6659)
RISC-V 优化
riscv 上 im2col gemm 与 winograd convolution 统一 elempack 优化 (Tencent#6740)
riscv convolution packed 优化 (Tencent#6731)
riscv gemm fp16s 支持 (@Xinyu302, Tencent#5311)
新增 riscv DeformableConv2D RVV 实现,相比 scalar 实现提速 12.94x ~ 20.16x (@chenglimin, Tencent#6540)
新增 RVV 1.0 Quantize layer 实现 (@justin Fung, Tencent#6636)
新增 RVV 1.0 Dequantize layer 实现 (@justin Fung, Tencent#6658)
新增 RVV 1.0 Requantize layer 实现 (@justin Fung, Tencent#6695)
新增 RVV threshold 实现 (@ihb2032, Tencent#6676)
新增 RVV shrink 实现 (@ihb2032, Tencent#6671)
新增 RVV softplus 实现 (@ihb2032, Tencent#6635)
新增 RVV power 实现 (@ihb2032, Tencent#6666)
新增 RVV log 实现 (@ihb2032, Tencent#6638)
新增 RVV exp 实现 (@ihb2032, Tencent#6637)
新增 RVV dropout fp16 实现 (@ihb2032, Tencent#6667)
修复 riscv 转 fp16 的隐式转换 warning (@bluemiao3, Tencent#6525)
MIPS / LoongArch
大规模 MIPS / LoongArch 优化,新增 absval / batchnorm / binaryop / bnll / clip / concat / conv1d packed 与 bf16s / conv 3x3 winograd 与 bf16s 与 int8 / im2col gemm 与 bf16s 与 int8 / gridsample / groupnorm / instancenorm / layernorm /
lrn / lstm / matmul / mha / sdpa / rmsnorm / rotaryembed / packing / padding / pooling / prelu / quantize / relu / reshape / scale / selu / mish / shufflechannel / sigmoid / slice / softmax 与 bf16s / swish / tanh 等大量优化实现
(Tencent#6662)
mips 平台支持 elu / erf / gelu / selu (@futz12, Tencent#6607)
修复
修复 SSE ShuffleChannel 最后通道处理逻辑越界写问题 (@junwha, Tencent#5735)
修复 x86 临时缓冲区对齐不当导致的 ASan 报错 (Tencent#6703)
修复 i386 上 x86 bf16 GEMM 打包顺序,使 dpbf16 正确消费 k/k+1 对,AVX512BF16 上 32-bit GEMM 不再失败;同时让 mha int8 测试输入采用与 GEMM int8 相同的安全随机化 (Tencent#6708)
修复 modelwriter 在空 per_channel_pad_data 时崩溃问题 (Tencent#6533)
修复 modelwriter 可选权重序列化 (@噜小噜芋团, Tencent#6726)
修复 fuse_padding_convolution 在不对称 padding 参数下被跳过 (@Mollyji, Tencent#6661)
修复 pnnx Conv2d 4 元素 padding 元组从 (left,right,top,bottom) 到 (height,width) 的归一化 (@Yeuvoir, Tencent#6694)
修复 pnnx 表达式 erf 应生成 Erf layer (@crafcat7, Tencent#6677)
修复 windows-arm 编译 (Tencent#6699)
修复 macos arm64 交叉编译时 x86_64 架构检测错误 (@YinHanke, Tencent#6730)
修复 NCNN_SIMPLEMATH 编译,并在 benchmark 改动时触发 CI (Tencent#6722)
修复 MSVC 上 NCNN_MALLOC_OVERREAD padding 缺失 (@ihb2032, Tencent#6583)
重构 BF16 转换逻辑,绕过 OHOS clang 在 aarch64 上的崩溃 (Tencent#6725)
针对 Clang 取消 -Ofast (其等价于 -O3 -ffast-math),避免版本不一致与未来废弃 (@zhuzeitou, Tencent#6520)
工具源码使用 snprintf 替代 sprintf (@evgeny Proydakov, Tencent#6554)
pnnx python 包将一处 bare except 替换为 except Exception (@Sense_wang, Tencent#6555)
PNNX
升级 pnnx 到 torch-2.10,附 GRU 等 pass_level2 适配 (Tencent#6592)
升级 pnnx 到 torch-2.11,更新 onnx 测试基线 (Tencent#6701)
pnnx 支持 npy 格式输入张量 (Tencent#6700)
pnnx 输出每层的 FLOPs 和 MemOps 统计 (Tencent#5836)
pnnx 将 prelu (num_parameters=1) 转换为 leakyrelu,使其能与 conv 融合 (@e8g(w7n), Tencent#6344)
pnnx 增加 TNN Flatten 算子 pass_level2 支持 (@yiming, Tencent#6513)
通用 / 构建
去除 vulkan/arm/riscv 上一些 layer 的虚继承 (Tencent#6590)
新增 cmake 选项使用系统提供的 pybind11 (@Integral, Tencent#6516)
工具与 datareader 路径下统一使用 snprintf (@evgeny Proydakov, Tencent#6554)
删除已过时的 msa.h gcc<8.5 workaround 文档 (Tencent#6734)
docs 修正 README 语法 (@4ek0, Tencent#6732)
README 集中合并多次更新 (Tencent#6739)
新增 Android 上通过 AHB 进行零拷贝输入的开发指南,含 ANDROID_API≥26、AImageReader_newWithUsage、每 AHB 指针 pipeline 缓存(约 5000x 启动开销减少)、Adreno 830 与 Mali-G925 跨厂商验证、ex.input(VkMat) 不会自动转格式的注意事项
(@securekim, Tencent#6733)
更新 convertmodel 网站使用文档 (@futz12, Tencent#6617)
docs/how-to-build 更新 RHEL/CentOS 编译依赖说明 (@bkmgit, Tencent#6692)
Benchmark
新增 Microsoft Azure 多种实例 benchmark 数据 (@kenji Mouri, Tencent#6552)
新增 Qualcomm Snapdragon X Elite benchmark 数据 (@Ratizux, Tencent#6535)
CI
修复 PR comment workflow (Tencent#6563)
release-python 不再使用 setup-python,避免污染 cibuildwheel 环境 (Tencent#6510)
更新 riscv ci toolchain 与 qemu,精简 workflow 117 行 (Tencent#6742)
增加 fp16s 测试,跳过 fp16p 在 cpu 上的测试 (Tencent#6724)
依赖更新
bump pypa/cibuildwheel 3.3.1 → 3.4.1
bump actions/github-script 8 → 9
bump GuillaumeFalourd/setup-windows10-sdk-action 2.4 → 2.5
bump docker/setup-qemu-action 3 → 4
bump actions/cache 4 → 5
bump softprops/action-gh-release 2 → 3
bump actions/upload-artifact 5 → 7
bump actions/download-artifact 7 → 8
bump codecov/codecov-action 5 → 6
New Contributors
@Integral made their first contribution in Tencent#6516
@yiming (leiyiming6) made their first contribution in Tencent#6513
@zhuzeitou made their first contribution in Tencent#6520
@bluemiao3 made their first contribution in Tencent#6525
@Ratizux made their first contribution in Tencent#6535
@chenglimin made their first contribution in Tencent#6540
@Sense_wang (haosenwang1018) made their first contribution in Tencent#6555
@NKID00 made their first contribution in Tencent#6615
@justin Fung made their first contribution in Tencent#6636
@crafcat7 (Lune) made their first contribution in Tencent#6657
@Edwardssss made their first contribution in Tencent#6682
@Roundaboutt (Layla) made th...
Release 20260112
blacklist slow emulated cooperative matrix on amd rdna2 (#6504)
Release 20260109
adapt more arm processor feature bits in newer windows sdk (#6498)
Release 20250915
disable experimental free threading wheel build
Release 20250912
fix android asset datareader scan to read from NULL-terminated string…
Release 20250822
Release 20250425
fix apple glslang package, drop glslang-default-resource-limits (#6022)
Release 20250425-2
do not install glslang when compiling shared library
Release 20241224-3
build for android 16k pagesize by default, update android api 21 (#5833)