Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/daily-til.lock.yml

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion .github/workflows/daily-til.md
Original file line number Diff line number Diff line change
Expand Up @@ -142,7 +142,7 @@ Read all five with a single `cat` each, then choose the topic. Never run shell s
- **Description**: one-sentence summary of the verdict/scope, 40-320 chars.
- **Body**: markdown tables for comparisons (the corpus uses them heavily) — every table cell padded with one space on each side of every pipe, like `| Model | Price |` (compact `|Model|Price|` fails lint); fenced code blocks with a language wherever commands or config appear; citations ONLY as `[label](https://...)` — never `[[url]]` wiki-links and never bare URLs, including in source tables; a blank line before and after every heading and every list; file ends with a newline; a `<` followed by a letter or digit in prose (`<50ms`, `<model>`) is JSX to MDX and fails the compile (so is a bare `{`) — write `under 50ms`, escape it as `\<50ms` / `\{`, or put it in backticks.
- **Sources must be reachable**: only domains on this workflow's network allowlist can be fetched (github.com, huggingface.co, openai.com, anthropic.com, ai.google.dev, blog.google, mistral.ai, deepseek.com, z.ai, qwencloud.com, python.org and their subdomains). Prefer candidates whose primary sources live there; if a candidate's key sources are blocked by the firewall, pick a different candidate rather than writing from memory.
- **Cite the paper that describes THIS release.** A model card often carries several `arxiv:` tags — the current technical report plus its predecessors. Picking the first or the oldest misattributes the work: PR #77 cited arXiv 2310.10688 (the 2023 original TimesFM paper) as TimesFM 3.0's reference, and PR #85 cited arXiv 2506.07900 as MiniCPM5-2B's technical report while the card also lists the newer 2602.09003. Read every `arxiv:` tag, prefer the newest one unless the card explicitly ties an older ID to this version, and if you cannot tell which describes this release, cite the model card and say the technical report was not identified.
- **Open every paper you cite and read its title before citing it.** An `arxiv:` tag on a model card is not a promise that the paper describes that model — cards routinely tag a predecessor's report, or an unrelated lab paper. Never infer a paper's subject from its ID, its position in the tag list, or its date. Fetch `https://arxiv.org/abs/<id>` and read the actual title: if it does not name this model and version, it is not this model's technical report. Say "no technical report is linked from the model card" rather than promote the closest-looking tag. Both previous failures came from guessing: PR #77 cited arXiv 2310.10688 as TimesFM 3.0's reference when it is the 2023 original TimesFM paper, and PR #85 cited arXiv 2506.07900 as MiniCPM5-2B's technical report when its title is "MiniCPM4: Ultra-Efficient LLMs on End Devices" — that card's other tag, 2602.09003, is a data-management paper, so neither was a match and picking "the newer one" would also have been wrong.
- **Prices come from the vendor's price sheet, never from arithmetic.** A relative claim ("one-tenth the price", "half the cost") is worthless without its baseline: read the sentence and name what it is cheaper *than* — vendors almost always mean their own previous model, not a competitor. Then fetch the vendor's pricing page (`docs.z.ai/guides/overview/pricing`, `qwencloud.com/models/<model>`, `openai.com/api/pricing`, `anthropic.com/pricing`) and quote the listed per-token numbers with that link. If no price is published, the cell reads "no official listing" — never multiply or divide some other model's price to produce a dollar figure, and never present a derived number as a price.
- **Never**: fabricated testing claims, "In today's fast-paced world" intros, unsupported superlatives, uncited numbers, emoji anywhere (headings, tables, lists — use "Yes"/"No" in comparison tables, plain words everywhere else).
- **Every contender must be real and fetched**: each row of a comparison table names a specific product you fetched a primary source for this run, with that source linked in the row or the section about it. Never invent placeholder contenders ("Generic X Server", "X Toolkit") to fill a table. If you cannot source at least three real contenders, write a single-product review of the one you can source instead of a comparison.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ GLM-5.3-Flash and Qwen3.8-Flash-Next represent two approaches to efficient large

| Model | Total Parameters | Active Parameters | Context Length | Price (per M tokens) | Best For |
| ------- | ------------------ | ------------------- | ---------------- | ---------------------- | ---------- |
| GLM-5.3-Flash | 320B | 18B | 300K tokens | $0.15 in / $0.50 out* | Multimodal agentic workflows |
| GLM-5.3-Flash | 320B | 18B | 1M tokens | $0.15 in / $0.50 out* | Multimodal agentic workflows |
| Qwen3.8-Flash-Next | 125B | 6B + 51B n-gram | 262K tokens (1M with YaRN) | No official listing† | Long-horizon reasoning & tool use |

\*List price from the [Z.ai pricing page](https://docs.z.ai/guides/overview/pricing) (cached input $0.03). A 50% launch discount ran through 2026-09-09.
Expand Down Expand Up @@ -144,7 +144,7 @@ Both models recommend specific settings for different modes:

### Context Management

- GLM-5.3-Flash uses explicit context management strategies for its 300K token evaluations
- GLM-5.3-Flash's window is 1M tokens (`max_position_embeddings` = 1,048,576); the 300K figure in its benchmark footnotes is the evaluation context for HLE, not the model limit
- Qwen3.8-Flash-Next natively supports 262K tokens and recommends YaRN for extension beyond that
- Both require careful output length allocation for agentic workflows (separate reasoning vs. final response limits)

Expand Down
9 changes: 7 additions & 2 deletions content/blog/glm-5-3-review-2026.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@
title: "GLM-5.3 Review: Z.ai's New Open Weights Model for Coding and Agentic Workflows"
description: "Comprehensive review of GLM-5.3, Z.ai's latest open weights Mixture-of-Experts model showing strong performance on coding benchmarks and cybersecurity tasks."
date: "2026-09-03"
updated: "2026-09-08"
author: "Sameer Khan"
tags: ["AI", "GLM", "Z.ai", "LLM", "Developer Tools", "Agentic AI"]
category: "AI"
Expand Down Expand Up @@ -29,7 +30,7 @@ GLM-5.3 was released by Z.ai in August 2026 as the latest iteration in their GLM
| AutomationBench (v1.0.6) | **48.2** | 26.2 | 43.2 | 39.8 | 41.0 | 45.8 |
| HLE w/ Tools | 62.5 | 54.7 | 60.0 | 56.2 | 57.9 | **64.5** |

**Context Length:** 300K tokens (with context management strategy)
**Context Length:** 1M tokens (`max_position_embeddings` = 1,048,576 in [config.json](https://huggingface.co/zai-org/GLM-5.3/blob/main/config.json))
**Reasoning Effort:** Configurable (low, high, max) - defaults to max
**License:** Other (GLM-5.3 license)
**Deployment:** Compatible with SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth, and Ascend NPU platforms
Expand Down Expand Up @@ -102,7 +103,7 @@ As an open weights model, GLM-5.3 can be run locally at no licensing cost (beyon
### Limitations

- The GLM-5.3 license contains restrictions that may affect certain use cases
- Context length of 300K requires careful management for optimal performance
- Long-context runs need careful management: Z.ai's own evaluations cap context per benchmark (300K for HLE, 400K for DeepSWE and Terminal-Bench 3.0, 1M for NL2Repo and ALE) rather than using the full window
- Some benchmarks show trailing performance compared to GPT-5.6 Sol
- Limited public API availability information in primary sources

Expand Down Expand Up @@ -135,3 +136,7 @@ For developers focused on coding agents, automation workflows, or cybersecurity

- GLM-5.3 model card and README: [GLM-5.3 model card and README](https://huggingface.co/zai-org/GLM-5.3) (accessed 2026-09-03)
- GLM-5.3 API metadata: [GLM-5.3 API metadata](https://huggingface.co/api/models/zai-org/GLM-5.3) (accessed 2026-09-03)

---

*Correction (2026-09-08): the original version listed the context length as 300K tokens. That figure is the evaluation context used for the HLE benchmark in the model card's footnotes, not the model's window; `config.json` gives `max_position_embeddings` = 1,048,576.*
9 changes: 7 additions & 2 deletions content/blog/minicpm5-2b-review-2026.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@
title: "MiniCPM5-2B Review: 2B Parameter Model Excels in Coding and Agent Tasks"
description: "Review of OpenBMB's MiniCPM5-2B, a dense 2B Transformer model achieving 2B-class SOTA performance with strong capabilities in coding, mathematics, and tool use."
date: "2026-09-08"
updated: "2026-09-08"
author: "Sameer Khan"
tags: ["AI", "LLM", "MiniCPM", "OpenBMB", "2B model"]
category: "AI"
Expand Down Expand Up @@ -70,7 +71,7 @@ MiniCPM5-2B is readily available through:
- Hugging Face Model Hub: `openbmb/MiniCPM5-2B`
- GitHub Repository: [OpenBMB/MiniCPM](https://github.com/OpenBMB/MiniCPM)
- Online Demo: Hugging Face Spaces
- Technical Report: arXiv:2506.07900
- Predecessor technical report: [MiniCPM4: Ultra-Efficient LLMs on End Devices](https://arxiv.org/abs/2506.07900) (arXiv:2506.07900, June 2025) — covers MiniCPM4, not this release; no MiniCPM5 technical report is linked from the model card

The model is part of the broader MiniCPM ecosystem, which includes various sizes and specialized variants to meet different deployment needs.

Expand All @@ -86,4 +87,8 @@ For teams working on resource-constrained applications or specialized developer

[2] MiniCPM5-2B README. Hugging Face. Accessed 2026-09-08. [https://huggingface.co/openbmb/MiniCPM5-2B/resolve/main/README.md](https://huggingface.co/openbmb/MiniCPM5-2B/resolve/main/README.md)

[3] MiniCPM Technical Report. arXiv. Accessed 2026-09-08. [https://arxiv.org/pdf/2506.07900](https://arxiv.org/pdf/2506.07900)
[3] MiniCPM4: Ultra-Efficient LLMs on End Devices. arXiv, June 2025. Accessed 2026-09-08. [https://arxiv.org/abs/2506.07900](https://arxiv.org/abs/2506.07900)

---

*Correction (2026-09-08): the original version cited arXiv:2506.07900 as MiniCPM5-2B's technical report. That paper is "MiniCPM4: Ultra-Efficient LLMs on End Devices" (June 2025) and describes the previous generation. The model card carries a second tag, arXiv:2602.09003, which is an unrelated data-management paper — neither is a MiniCPM5 report, and the card links none.*
Loading