Skip to content

Repository files navigation

Chinese Name Meaning Explorer | 中文姓名解析计划

English | 中文


English

A Vue 3 application that analyzes Chinese names through character definitions, cultural context, and local AI analysis to provide a deep understanding of each character's meaning.

Features

  • Hanzi-Specialized Analysis: Optimized for 2-4 character Chinese names in the current U+4E00-U+9FA5 validator range, with automatic surname/given-name segmentation.
  • Deep Dictionary Integration: Uses Chinese-first dictionary data with a reviewed CC-CEDICT supplement overlay for missing or damaged definitions.
  • Context-Aware Readings: Applies single and compound surname pronunciations only after surname segmentation, including polyphonic surnames such as 乐, 翟, 华, and 覃.
  • Cultural Context: Includes Five Elements, literary references, gender bias, and naming connotations.
  • Optional Guangyun View: An opt-in character-card section displays source-separated Guangyun fanqie, rhyme metadata, and historical glosses without treating them as modern or naming meanings.
  • Local AI Model (ONNX): Uses a custom-trained 10-label classifier (Scholarly, Heroic, Serene, etc.) with WebGPU hardware acceleration.
  • Layered Local Inference: ONNX predicts the 10 tone labels. Tauri first tries a downloaded native Qwen2.5 GGUF; a missing native model permits Ollama, while native timeout, runtime, and quality failures are reported explicitly.
  • Grounded AI Narratives: Qwen receives only parsed character facts and explicit fact boundaries. Unsupported biography, history, pinyin leakage, prompt repetition, or invalid output length triggers a corrective retry rather than a deterministic result presented as Qwen output.
  • Curated Literary Context: Source-specific character allusions may enrich the analysis, while famous-person matching is deliberately excluded from automatic name interpretation.
  • Controlled Inference Lifecycle: Worker, Ollama, and native requests have bounded timeouts, cancellation, and cleanup. Native Qwen attempts at most three times; ONNX Worker inference remains limited to two attempts.
  • Privacy & History: 100% local processing; history is stored in browser localStorage.
  • Open Feedback Loop: Integrated GitHub feedback system with automated environment diagnostics.

Tech Stack & Architecture

  • Frontend: Vue 3 (Composition API)
  • Engine: localInference.ts orchestrates an ONNX Runtime Web worker, native Tauri commands, local Ollama, and deterministic fallback.
  • Acceleration: Prioritizes WebGPU with a stable WebAssembly fallback.
  • Desktop: Packaged as a native Windows .exe via Tauri.
  • Native LLM: Rust llama-cpp-2 loads a Qwen2.5 0.5B Q4_K_M GGUF. The download dialog pre-fills %LOCALAPPDATA%\Chinese Name Meaning Explorer\models and accepts a persistent custom absolute directory such as a D drive folder.
  • Model Integrity: The 491,400,032-byte download is pinned to a Hugging Face revision and verified by HTTP status, size, and SHA-256 before atomic installation.
  • Resilient Download: The desktop downloader falls back to a mirror when the primary Hugging Face endpoint is unavailable without relaxing size or SHA-256 verification.

中文

一个基于 Vue 3 的中文姓名解析应用,通过汉字字义、文化背景及本地 AI 分析,深度解读每一个汉字背后的意义。

功能特性

  • 汉字特化解析:专门针对当前校验范围 U+4E00-U+9FA5 内的 2-4 位中文姓名进行优化,自动识别单姓、常见复姓与名字。
  • 姓名上下文读音:通用字典保留普通字音,单字姓数据库与特殊复姓读音表在姓氏位置覆盖多音字读音,例如 乐 yuè、翟 zhái、单于 chányú。
  • 深度字义解析:整理自新华字典等来源,提供中文释义;运行时会过滤空白或纯标点残片,避免把无意义字典内容显示给用户。
  • 文化背景关联:集成五行属性、典故出处、性别倾向及命名寓意。
  • 本地 AI 模型 (ONNX):使用本地训练的 10 标签分类器(如书卷、豪迈、灵动等),通过 WebGPU 硬件加速进行“意境”实时分析。
  • 分层本地推理:ONNX 负责 10 类意境标签;Tauri 桌面端优先使用按需下载的 Qwen2.5 GGUF。原生不可用或非超时失败时尝试 Ollama;原生超时则直接回退到确定性文本。
  • 事实约束叙述:Qwen 仅接收解析后的汉字事实和明确事实边界;包含无来源人物传记、历史引申、拼音泄漏、提示复述或长度不足的输出会被拒绝并回退到确定性本地文本。
  • 可核验典故:经过核验的汉字典故可用于丰富意境分析;历史人物同名匹配不会自动进入姓名解释,避免把用户误识别为名人。
  • 可控推理生命周期:Worker、Ollama 和原生推理均具备有界超时、取消、资源清理和最多两次尝试。
  • 隐私与历史:所有数据本地加载,历史记录存储于浏览器 localStorage,不上传任何隐私。
  • 反馈闭环:内置 GitHub 反馈入口,自动收集基础诊断信息。

技术架构

  • 前端框架:Vue 3 (Composition API)
  • 推理引擎:localInference.ts 统一调度 ONNX Worker、Tauri 原生命令、本地 Ollama 和确定性回退。
  • 硬件加速:优先尝试 WebGPU,稳健回退至 WebAssembly。
  • 桌面支持:通过 Tauri 提供 Windows .exe 原生包支持。
  • 原生大模型:Rust llama-cpp-2 加载 Qwen2.5 0.5B Q4_K_M GGUF。下载弹窗默认预填 %LOCALAPPDATA%\Chinese Name Meaning Explorer\models,也可持久化使用 D 盘等自定义绝对目录。
  • 模型完整性:491,400,032 字节的下载固定到 Hugging Face revision,并在原子安装前校验 HTTP 状态、大小和 SHA-256。
  • 下载容错:官方 Hugging Face 端点不可用时自动切换镜像,同时保持大小和 SHA-256 完整性校验不变。

Project Structure | 项目结构

my-vue-app/
  public/
    data/
      chars.json      # Chinese-first dictionary (Xinhua core data)
      surnames.json   # Single-character surname-specific readings
    models/
      classifier.onnx # Custom-trained 16-dim feature -> 10-class classifier
      manifest.json   # Model version and label mapping
  src/
    App.vue           # Core UI logic
    services/
      localInference.ts       # AI Orchestration
      nameAnalyzer.ts         # Dict & Segmentation engine
    data/
      compoundSurnamePinyin.json # Contextual readings that cannot be inferred safely
    workers/
      localInference.worker.ts # ONNX Inference worker
  src-tauri/
    src/main.rs                # Secure model download and Tauri commands
    src/native_llm.rs          # llama.cpp-backed GGUF generation

ML Context | 本地训练与模型

If you wish to retrain the model, use my-vue-app/train_model.py (requires torch/onnx):

  • Labels: Scholarly, Grand, Heroic, Serene, Classical, Unique, Dynamic, Persistent, Nature, Deep.
  • Feature Engineering: 16-dimensional hybrid vector including 4 semantic category scores.

如需重新训练模型,请使用 my-vue-app/train_model.py:

  • 标签体系:书卷、宏伟、豪迈、恬静、典雅、新颖、灵动、坚毅、自然、深邃。
  • 特征工程:16 维混合向量,注入了 4 类语义词谱得分。

Development | 开发指南

cd my-vue-app
npm ci
npm run dev

Windows Packaging | 打包发布

  1. Install locked dependencies in my-vue-app: npm ci.
  2. Ensure Rust stable, MSVC, Visual Studio C++ build tools, CMake, and LLVM/libclang are installed on Windows.
  3. Run npm run tauri:build.
  4. Output: my-vue-app/src-tauri/target/release/bundle/.

The installer contains the downloader, not the GGUF weights. On first desktop launch, users may download approximately 491 MB to the pre-filled default directory or a custom absolute directory. When the model is missing, the download dialog warns systems reporting less than 6GB RAM; it does not block download or later native inference.

Verification | 验证步骤

cd my-vue-app
npm audit --omit=dev
npm audit
npm run test:features
npm run test:unit -- --run
npm run type-check
npm run lint:check
npm run test:onnx
npm run build
cd src-tauri
cargo fmt --check
cargo check --locked
cargo test --locked

test:onnx creates an ONNX Runtime Web WASM session, executes public/models/classifier.onnx with a real tensor, and validates the output contract and finite values. lint:check is read-only and scans only project-owned source, scripts, and configuration; generated ONNX Runtime files under public/ are excluded.

The unit suite also loads the production chars.json and surnames.json files to verify polyphonic single surnames, explicit compound-surname readings, Pinyin normalization, and punctuation-only definition cleanup. Keep surname-specific readings out of the general character dictionary: they are applied only after segmentation identifies surname context.

The tag-triggered Windows workflow runs version validation, feature tests, unit tests, TypeScript checking, read-only lint, the real ONNX smoke test, and locked Rust checks before tauri-apps/tauri-action@v0 can package a release. The committed npm dependency graph currently reports zero vulnerabilities through both npm audit --omit=dev and npm audit.

Data Sources & License | 数据来源与证书

  • Dictionary supplements: CC-CEDICT, official 2026-08-02 release, CC BY-SA 4.0. Reviewed Simplified Chinese translations preserve their source lines and English glosses in a reproducible overlay.
  • Bulk dictionary corroboration: Unicode Unihan 17.0.0 kDefinition, Unicode License V3. Unihan glosses remain separate evidence and are never published as unreviewed machine translations.
  • Pure-Chinese definition supplements: Chinese Wiktionary 2026-07-01 fixed dump, CC BY-SA 4.0. Each published definition links to an immutable page revision and records the selected source text and Simplified Chinese transformation; unreviewed parser output remains isolated from production data.
  • Legacy consolidated definitions remain subject to provenance review; new supplementation does not use scraped Xinhua or Kangxi mirrors with an unclear redistribution chain.
  • License: CC BY-SA 4.0.

About

Privacy-first Chinese name explorer with surname-aware parsing, cultural insights, local ONNX inference, and optional Tauri/Qwen desktop AI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages