Skip to content

Latest commit

 

History

History
196 lines (155 loc) · 7.11 KB

File metadata and controls

196 lines (155 loc) · 7.11 KB
contract-id hackylens.ai-model-package
owner ai-runtime
version 1.0.0
stability experimental
schema-major 1

AI models

HackyLens separates reusable K210 model mechanics from feature behavior.

Firmware boundary

core/ai_model_types.h describes:

  • the exact KModel v3 header and first-layer DMA contract;
  • input and output tensor shape, element type, layout, and normalization;
  • an optional expected model CRC32;
  • labels and the post-processing category;
  • SD paths, size limits, and unload timeout.

storage/ai_model_storage.* mounts FAT32, reads a model into a 256-byte-aligned allocation, calculates CRC32, and parses the optional 256-byte manifest.hkai. services/ai_model_runtime.* owns the single KPU lease, validates the descriptor, manifest, model identity, and output sizes, runs asynchronous inference, and completes deferred unload. The shared layer does not know about cameras, labels, boxes, or a particular detector.

FACE and OBJECT share services/camera_ai_input.*. It arms the planar DVP output only at a camera frame boundary, freezes it on that frame's finish, and keeps the input immutable until KPU completion. OBJECT decodes the completed float output on core 0: the K210 cores have non-coherent caches, so passing the SDK's CPU-dequantized tensor to core 1 would require an unnecessary extra coherency protocol.

Included consumers

FACE DETECT continues to load:

/hackylens.kmodels/detect.kmodel

The included FACE model is the native KModel v3 from Kendryte's kendryte-standalone-demo/face_detect example, pinned at commit e89c35465fadb5524c892d2a1c7a76dc76e219ed:

upstream path   face_detect/detect.kmodel
model bytes     388,776
SHA-256         916e679defa91ad76f9feed18b6b37d26328ec9a2c0c8ab0d1ca5983e105b7c0
CRC32           40429f51
input           U8 CHW  [1, 3, 240, 320]
output          F32 CHW [1, 30, 15, 20]
layers          24

The upstream README says the model was generated by nncase, but does not name the training dataset or original training source. The repository says demo licenses are specified separately, while the pinned face_detect directory does not contain a separate license for the code or pretrained weights. These unknowns are retained explicitly in sdcard/hackylens.kmodels/detect.UPSTREAM.txt.

This file must not be replaced with the 386,608-byte detect.kmodel extracted from the original HUSKYLENS firmware package. That asset stores 32-bit words in flash byte order and uses a different/private container. Large weight regions match the standalone model, but the original asset is not directly loadable by HackyLens's native KModel-v3 runtime.

OBJECT DETECT requires:

/hackylens.kmodels/object20/model.kmodel
/hackylens.kmodels/object20/manifest.hkai
/hackylens.kmodels/object20/labels.txt

Its pinned model is the official Kendryte nncase 20classes_yolo example at commit f92c085ec6355ae04258df5b76ec7570f784129d:

model bytes     1,352,588
SHA-256         33219de6ffa0b24b8c41a82d09888070c72be68356ca457f3b722d25fdff95ab
CRC32           107d903c
input           U8 CHW  [1, 3, 240, 320]
output          F32 CHW [1, 125, 7, 10]
classes         Pascal VOC20
anchors         1.08,1.19  3.42,4.41  6.63,11.38  9.42,5.11  16.62,10.52
defaults        confidence 0.50, class-aware NMS 0.20

The repository containing the example is Apache-2.0 licensed. Its example directory does not provide a separate provenance statement for the pretrained weights, which is recorded explicitly rather than silently assigning one.

SD staging

The tracked sdcard/ directory mirrors the card root. Copy its contents to a FAT32 card without adding another directory level:

sdcard/
└── hackylens.kmodels/
    ├── detect.kmodel
    ├── detect.UPSTREAM.txt
    └── object20/
        ├── model.kmodel
        ├── manifest.hkai
        ├── manifest.json
        └── labels.txt

The object package can be reproduced from its pinned URL, byte size, and SHA-256:

python tools\ai_model.py fetch `
  --spec models\object_detect_voc20.json `
  --out-dir sdcard\hackylens.kmodels\object20

python tools\ai_model.py verify sdcard\hackylens.kmodels\object20 `
  --spec models\object_detect_voc20.json

manifest.hkai is the compact record consumed by firmware. manifest.json keeps the full audit data, including SHA-256, anchors, defaults, and upstream commit. OBJECT also pins the expected CRC32 in its descriptor, so another same-shaped model cannot accidentally use the fixed VOC20 decoder.

Model lab

tools/ai_model.py provides:

  • inspect for native KModel v3 and recovered original model formats;
  • fetch for a spec-pinned upstream model plus SD packaging;
  • package for a locally supplied native KModel v3;
  • verify for binary/JSON manifest, CRC, SHA, output, and label checks;
  • bootstrap and convert for the pinned legacy compiler.

The locked compiler is Kendryte nncase v0.1.0-rc5, which emits the KModel v3 container supported by the current K210 runtime. Its real CLI is flat:

ncc -i tflite -o k210model --dataset calibration-images \
    --inference-type uint8 --channelwise-output input.tflite output.kmodel

The wrapper supplies --channelwise-output automatically for a YOLO2 spec:

python tools/ai_model.py convert \
  --source detector.tflite \
  --format tflite \
  --calibration calibration-images \
  --spec models/spec.example.json \
  --out-dir out/detector

Rc5 directly supports TFLite and Caffe frontends. It does not accept ONNX, modern compile subcommands, -t k210, or modern allocator/calibration switches. ONNX needs a separately validated ONNX-to-TFLite step. Moving to nncase 0.2+ would produce KModel v4 and requires a different device runtime.

Conversion success is only the first gate. Every new model still needs framework-versus-KModel comparison, operator review, KPU/main-memory measurement, device latency, and accuracy tests.

Original firmware findings

The original package's unpacked/detect.kmodel is a 386,608-byte flash-oriented/private representation with SHA-256 af5ac498c67fb65bd4a49553c27546c98d4f53e503519eb9790d4d8a067e0fae. After reversing each stored 32-bit word it shares large exact network payload regions with the standalone FACE model, but its header and serialized layer metadata are different. It is retained only as a reverse-engineering artifact, not as the SD-loadable FACE model.

unpacked/object_detect.bin is not a data store. It is the original HUSKYLENS custom, 32-bit word-swapped legacy KPU task container:

SHA-256         c6b34d98aafa94f6a5fa7c3429b0e6cf384e5f5d4769793da12d7e0267e86e26
input           U8 CHW [1, 3, 240, 320]
output          U8 CHW [1, 125, 7, 10]
layers          16

It uses a private loader and output dequantization path and cannot be passed to the current native KModel v3 runtime without a legacy backend or verified conversion. tools/ai_model.py inspect reports its known contract only when the exact SHA-256 matches.

unpacked/mobilenetv1_1.0.kmodel is instead a 224x224 classifier with one 1,000-float output. It does not produce bounding boxes and is not the 20-class OBJECT DETECT model.