Skip to content

🗣 Daily social checklist — 2026-09-23 #683

Description

@github-actions

🗣 Daily social post

benchmark records collected

Ready-to-post material for today. Rephrase per platform; the two sentences below are the factual core.

Benchmark update: Today's radar surfaced 510 items across Hugging Face, Crossref, arXiv, GitHub, Zenodo, Hugging Face Papers, OpenAlex, Kaggle Dataset, First-party feed, GitHub Release, GitHub Organization; top signal: A MARL Simulation Benchmark and Systematic Evaluation for Multi-UAV Cooperative 3D Voxel Coverage (Crossref, 69/100).

Repo change: 1 commit in the last 24 hours across registry data.

发布文案示例 (copy-paste for today's post):

发前提醒:涉及条目数、来源数、star、日期或榜单名次时,先用当天数据替换<最新数据>。标题优先说“免费使用”“一键导出”“开源”,少写“我做了什么”。
要多放软广。。少放硬的。

  • 在别人的微信群,需要cue 他们本地的bench。

  • 发微信,群发给朋友们。一人一个唾沫星子都有10k+ stars

  • 软广
    有没有家人们做个“说人话” bench?
    最近Claudish逼疯了 哈哈
    衡量LLM说人话程度: or 看看全网11923个 benchmark/eval/dataset消息
    仓库在 github.com/ktwu01/benchmark-radar,有空可以看看

  • Benchmark Radar:为 Agent 提供第一手消息
    做agent benchmark 的时候,找 related work 渠道太多,很难找全,找不全的话,做了研究又容易被人截胡。
    所以我们做了个一键生成 LaTeX 版本 related work 的 skill
    最近从github / huggingface 等37个来源,收集到了10k+ benchmarks/eval records
    欢迎试用和提意见 https://benchmark-radar.org/cli
    P.S. repo 在 https://github.com/ktwu01/benchmark-radar,已经 180+ stars

  • B站:每天整理 Agent benchmark
    Agent benchmark 每天都能出好几篇。最近看得比较多,我就做了
    Daily Agent Benchmarks:每天整理 GitHub arXiv 等等37个来源的新发布的 Agent 评测,
    每天整理新发布的 Agent 评测,支持按领域筛选、关键词搜索和引用数排序。
    已经有 11923+ 条数据,180+ stars 了,欢迎大家去看看,最近大家都在关注什么~

  • Agent benchmark 还在测“能不能把事做完”

    一个反常识的发现:Agent benchmarks 看起来已经在评测智能体前沿,
    实际上大多还在测“能不能按要求把事情做完”。

    按 Hejia Geng 的 L0–L5 分级,当前主体仍是 L2 执行和 L3 复现;
    L4 再发现刚开始出现。L5 要求 Agent 产出 benchmark 创建时未知的知识
    或方法,本周仍未发现新案例。

    源码和分析:Agent Benchmark 发展到什么阶段:L0-L5 能力层级与 SLIM 质量框架 #150

  • 免费一键导出全网 benchmark 数据

    免费开放,一键导出 <最新条目数> 条 benchmark、eval 和 dataset 数据。
    Benchmark Radar 每天从 <最新来源数> 个来源更新。如果你在找 related work、
    给自己的 Agent 选 eval,或者想跟进最新评测,可以直接搜,也可以订阅 RSS。

    项目:https://github.com/ktwu01/benchmark-radar

  • Free, open-source benchmark export

    Free to use and open source: export <latest record count> benchmark,
    evaluation, and dataset records in one step. Benchmark Radar updates daily
    from <latest source count> public sources and includes search, RSS, and an
    offline CLI.

    Use it to find related work, shortlist an eval for your agent, or follow new
    evaluation releases: https://github.com/ktwu01/benchmark-radar

  • 做 benchmark 前先查 related work

    做 Agent benchmark、找 related work 时,信息散在论文、仓库和榜单里,
    很容易漏掉,做完才发现被人截胡了。Benchmark Radar 每天收集一手来源,
    网页和 CLI 都能搜。先查相近工作,再决定自己的 benchmark 该补哪块空白。

    CLI 和数据:https://github.com/ktwu01/benchmark-radar

  • Frontier labs 到底在报哪些 benchmarks

    I checked 30 frontier model and system cards from 10 organizations and found
    79 distinct benchmarks. GPQA Diamond appeared in 23 of 30 documents,
    SWE-bench Verified in 18, and LiveCodeBench in 15.

    The count is document-level, not score-row-level. I excluded cross-vendor
    score comparisons because prompting, scaffolding, pass@k, tools, and
    evaluator settings often differ. The dataset and method are open for review:
    https://github.com/ktwu01/benchmark-radar

    Historical launch example from August 2026; re-check every number before reuse.

  • Agent 做了有害操作,不等于它在隐藏操作

    最近有前沿 AI 实验室公开 Agent 自主越界的证据,相关工作很快被做成了
    benchmark。评测同时记录系统真正看到的 action,以及模型自述的 action,
    再比较两份日志。

    结果里有模型没有输出 action log,有模型卡在 reasoning loop,也有模型执行了
    harmful action,但如实记录了行为。“是否有害”和“是否隐藏”是两个不同的
    安全问题。

    今日来源和原始证据:<当天 Benchmark Radar 深链接>

  • 只分享一个具体榜单,例如 AIME

    想看 AIME 在不同模型报告里的采用情况和可比成绩,不用先看项目介绍,
    直接打开这个榜单:
    https://benchmark-radar.org/leaderboard/?lfrontier=llm-stats-aime-2025

  • 小红书:新 benchmark 的热度该怎么看

    Benchmark Radar 的热度主要看三类公开数据:Hugging Face Paper 点赞、
    GitHub Stars 和 Hugging Face 数据集下载量。

    刚发布时,stars 和下载量通常还没积累起来,所以短期更看 HF Paper 点赞。
    时间拉长到 30 天或 90 天后,stars 和下载量更能说明社区是否持续关注、
    是否真的在使用。页面也会说明缺失数据和不同时间窗口的处理方法。

    #ai #benchmark

  • KOL 软推广:先问一个真问题

    有没有人做过“说人话” benchmark?最近有些模型的表达越来越像模板,
    我很想看一个专门衡量自然表达的评测。

    如果你也在设计新 benchmark,或者正在找 related work,可以用
    Benchmark Radar 搜公开论文、仓库、数据集和榜单:
    https://github.com/ktwu01/benchmark-radar

  • Cold email:你的 benchmark 进入了今日 Top 10

    Subject: Your new benchmark is in today's Benchmark Radar Top 10

    Hey! We found <benchmark name> in today's Benchmark Radar Top 10.
    Interested in using our CLI to find related works from 11923+ sources?

    CLI: https://benchmark-radar.org/cli
    Code and data: https://github.com/ktwu01/benchmark-radar

  • 可视化:网上在热议什么,最后沉淀成了什么

    上半张图看每年研究社区在讨论什么,下半张图看哪些 benchmark 最后进入了
    前沿模型报告。两组数据要分开:研究热度来自 discovery corpus,采用情况来自
    model-card reporting。这样才能看出一个话题从讨论走向评测基础设施的过程。

  • 我们给全网 11923+ AI benchmarks做了个CLI

    中文版

    做了个 Benchmark Radar。

    它每天收集新发布和新更新的 benchmark、评测、数据集与数据质量工作,
    帮你快速判断今天最值得先看什么。Today 页会优先展示真正的新发布,
    再看新鲜度和优先级分数。

    当前的优先级分数由四部分组成:相关性占 35%,社区采用信号占 25%,
    证据强度占 20%,新鲜度占 20%。相关性看标题和来源描述是否确实在讲
    benchmark 或评测;证据强度看有没有论文、代码、数据、作者和交叉链接;
    新鲜度则结合发布时间,以及这是首次发布还是普通更新。

    社区采用信号包括 GitHub Stars、论文引用、下载量和点赞数,但不会直接把这些数字相加。
    系统会按不同指标分别做对数归一化,再采用其中最强的一项。下载和点赞容易受到
    自动流量影响,因此分数设有上限,不能仅靠很高的下载量压过获得大量 Stars 或引用的项目。

    这个分数只用于信息筛选,回答“接下来值得点开哪个”,不代表 benchmark 的科学质量。
    Hacker News 等注意力信号会单独展示,不参与证据排名。LLM 只负责根据当天已采集、
    可引用的证据生成每日简报,不参与打分,也不会预测一个项目未来能有多火。

    English

    I built Benchmark Radar.

    It tracks newly released and updated benchmarks, evaluations, datasets,
    and data-quality work, helping you decide what is worth opening first.
    The Today page prioritizes new releases, followed by recency and priority
    score.

    The priority score has four components: relevance (35%), adoption signals
    (25%), evidence strength (20%), and recency (20%). Relevance checks whether
    the title and source description are genuinely about benchmarks or
    evaluations. Evidence strength considers papers, code, data, authors, and
    cross-links. Recency accounts for both timing and event type, distinguishing
    new releases from routine updates.

    Adoption signals include GitHub stars, citations, downloads, and likes.
    These raw counts are not added together. Each metric is normalized on its
    own logarithmic scale, and the strongest available signal is used. Downloads
    and likes are capped because they can accumulate through automated activity;
    high download counts alone should not outrank projects with substantial stars
    or citations.

    This score is for triage—answering “What should I open next?”—not judging
    scientific quality. Attention signals such as Hacker News activity are
    displayed separately and do not affect the evidence ranking. The LLM only
    produces a daily briefing from cited, collected evidence. It does not score
    records or predict future popularity.

Posting checklist - tick a channel after today's post is sent there:

Daily targets:

  • 众筹-social:请 benchmark 群友转发到自己的社群
  • X / Twitter
  • Benchmark Radar 讨论群
  • GitHub Blog https://ktwu01.github.io/
  • WeChat Moment
  • Science Intelligence 实名讨论群
  • Hacker News https://news.ycombinator.com/submit
  • 灵台AI
  • 小红书
  • 知乎
  • 即刻
  • Agent Infra交流群 (Substrate)
  • HLE V2 | Core
  • Cold email:arXiv / OpenReview benchmark 作者
  • 新微信群或 Discord 群
  • KOL:逛逛 GitHub 等相关账号
  • Bilibili / B站
  • 微信公众号
  • V2EX
  • YouTube
  • 抖音 / Douyin
  • TikTok
  • Instagram
  • Threads
  • Facebook
  • Bluesky
  • Mastodon
  • Telegram
  • Medium
  • DEV Community
  • Hashnode
  • Substack Notes
  • Indie Hackers
  • Product Hunt
  • DevHunt
  • AgentHunter
  • BetaList
  • Peerlist
  • AppSumo
  • Lobsters
  • GitHub Discussions
  • GitHub Awesome lists / resource-list PRs
  • Papers with Code discussions

Weekly (every 7 days):

Monthly:

  • LinkedIn
  • TB
  • TB Sci
  • X Ads
  • Google Ads
  • Meta / Facebook Ads
  • Bing Ads

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    automatedCreated by automationtractionGrowth, promotion, distribution, and audience-building work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions