Support: fuel the next build —
High-Throughput Synthetic & Pretraining Dataset Sifter Powered by TypeSafe AI (Jev)
Beta: v0.1.1 is experimental. Expect rough edges. Please contribute by opening an issue or PR.
Live: jev-curate.vercel.app (measured 24.0 rows/sec single-node on local mock bench examples/bench_mock.rs, 1,500+ cluster target)
Why jev-curate • Quickstart • CLI Reference • Python API • Architecture • Non-Goals • Ecosystem
OpenJEV support: Jev is built by TypeSafe. This fork keeps TypeSafe as the default and adds optional support for OpenJEV, a free community gateway to the same Jev model — set
OPENJEV_API_KEY(orJEV_PROVIDER=openjev) to use it. Original project: https://github.com/AkashPriyadarshii/jev-curate by @AkashPriyadarshii.
Cleaning 10M to 1B rows of synthetic reasoning data, instruction tuning pairs, or web-scraped corpora is an economic and technical nightmare:
- Generative LLMs are too slow and expensive: Running Claude 3.5 Sonnet or GPT-4o to judge synthetic rows costs $15,000–$50,000 per billion tokens and crawls at a painful 30–50 rows/sec.
- Regex heuristics are blind to reasoning flaws: Keyword and regex filters can check syntax, but fail to detect circular reasoning, hallucinated derivation steps, or robotic sycophancy.
- Context rot from uncompressed inputs: Naively feeding raw data into LLMs causes decision accuracy to crater while burning money on boilerplate text.
jev-curate solves this by piping Apache Arrow and Parquet streams through TypeSafe AI's Jev model (jev-1.13.0):
- Single-round-trip rubrics per row: Evaluates all rubric questions concurrently in one HTTP request per row via Jev's speculative parallel fan-out with zero per-question round-trips. Single-node throughput is bounded by TypeSafe's 1,200 req/min (20 rows/sec) limit; horizontal scaling across worker nodes targets 1,500+ rows/sec cluster throughput.
- ~$4.20 per 100M tokens: Jev bills $0.042/Mtok for input, zero for output. TypeSafe benchmarks System One workflows 444.6x cheaper and 193.6x faster than generative LLMs (source).
- Mathematical calibration: Receives calibrated probabilities (
Noul), ordinal rubrics (Scoreon a 0 to 4 scale, sent as an ordered list), and categorical choices (Choice), eliminating generative text slop. - Zero Rewriting: Emits clean records verbatim without rewriting or altering mathematical formulas.
# Install via Cargo
cargo install jev-curate
# Set your TypeSafe AI key (default provider)
export TYPESAFE_API_KEY="your-api-key"
# Or use OpenJEV (free community gateway to the same Jev model)
export OPENJEV_API_KEY="your-openjev-key"
# Filter a Parquet dataset using the math reasoning preset:
jev-curate filter train.parquet \
--preset reasoning-math \
--out ./output/ \
--concurrency 32
# Explicitly select OpenJEV without setting JEV_PROVIDER:
jev-curate filter train.parquet --provider openjev --concurrency 32PyO3 bindings (build from source with the python cargo feature):
git clone https://github.com/AkashPriyadarshii/jev-curate
cd jev-curate
maturin develop # python feature auto-enabled via pyproject.tomlfrom jev_curate import PyJevCurator
curator = PyJevCurator(api_key="your-api-key", preset="reasoning-math")PyJevCurator currently wraps the same filter pipeline as the CLI (see src/filter.rs); constructor-only for now. Use the CLI for row-level sifting.
jev-curate filter [OPTIONS] <INPUT_PATH>
| Flag | Default | Description |
|---|---|---|
<INPUT_PATH> |
Required | Path to input .parquet or .jsonl file. |
-p, --preset |
reasoning-math |
Pre-built rubric (reasoning-math, anti-sycophancy, code-correctness). |
-o, --out |
./curated/ |
Destination folder for clean.jsonl and rejected.jsonl. |
-c, --concurrency |
32 |
Worker concurrency (adaptive token bucket prevents 429 rate limits). |
--dry-run |
false |
Offline evaluation simulation with host pre-filtering and zero API calls (no TYPESAFE_API_KEY needed). |
--endpoint |
None | Custom API endpoint URL for offline mock testing (or set TYPESAFE_ENDPOINT). |
--provider |
auto | Jev provider: typesafe (default) or openjev (or set JEV_PROVIDER). |
| Preset | Primitives Evaluated | Target Problem Solved |
|---|---|---|
reasoning-math |
has_circular_reasoning (Noul)reasoning_depth (Score 0-4, keep 2.0 or higher) |
Drops ungrounded math derivations and repetitive circular proofs. |
anti-sycophancy |
is_sycophantic (Noul)has_ai_disclaimer (Noul) |
Eliminates "As an AI...", ungrounded flattery, and conversational filler. |
code-correctness |
has_stub_placeholders (Noul)code_quality (Score 0-4, keep 2.0 or higher) |
Drops incomplete code blocks and unrunnable pseudo-code mocks. |
# Rust core
cargo build --release
cargo test
# Python bindings (via maturin)
maturin develop
pytestAll tests run against an in-process mock server with zero live API credits in CI.
jev-curate/
├── Cargo.toml # Rust core manifest (arrow, parquet, tokio, pyo3, clap, reqwest, serde)
├── pyproject.toml # Maturin Python package manifest
├── src/
│ ├── lib.rs # PyO3 module bindings & crate entry
│ ├── main.rs # Standalone CLI binary entrypoint
│ ├── client.rs # TypeSafe AI HTTP client (speculative fan-out)
│ ├── filter.rs # Host-side sanity pruning & Jev pipeline
│ ├── parquet_io.rs # Streaming Parquet/Arrow reader and writer
│ ├── rate_limiter.rs # Adaptive token-bucket with auto 429 backoff
│ └── presets.rs # Pre-built post-training evaluation rubrics
└── tests/
└── mock_test.rs # In-process mock tests via typesafe-rs-mock (100% offline)
- Not a Generative Re-writer:
jev-curatenever paraphrases or re-generates text. Data is kept 100% verbatim. - Not a Heavy Local Vector DB: No embeddings, no vector indices, zero PyTorch/CUDA runtime requirements.
- Not a Generic Web Scraper: Tailored strictly for structured datasets (Parquet, Arrow, JSONL).
- X / Twitter • Threads • Instagram • Reddit
Built with high-performance Rust for the TypeSafe AI System One (Jev) ecosystem.
Keywords: TypeSafe AI, Jev, api.typesafe.ai, System One, Choice, Score, Noul, dataset curation, synthetic data filtering, pretraining datasets, Parquet streaming, arrow, rust.