Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions DEVELOPMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -150,7 +150,7 @@ def check_permission(state, cls, name, args): # 决策 +

```bash
python agent.py --selfcheck # 零依赖、零网络、零 key。改完先跑这个
python -m pytest tests/ -q # 337 条。CI 跑 ubuntu/windows × 3.10/3.13
python -m pytest tests/ -q # 343 条。CI 跑 ubuntu/windows × 3.10/3.13
```

`--selfcheck` 覆盖 4 个工具的读/写/改、`edit_file` 的"找不到 / 不唯一"两种报错、
Expand Down Expand Up @@ -190,7 +190,7 @@ set TALOS_MAX_STEPS=12
| **行为类** | **模型读到守卫那条消息之后干什么** | ❌ 必须用你真在用的那个 |

「拒绝之后会不会改用 `python -c` 绕路」是那个具体模型的性格,换了模型不算数。
而**行为类正是 live 测试唯一还值钱的部分** —— 管道通不通,337 条判据已经免费覆盖了。
而**行为类正是 live 测试唯一还值钱的部分** —— 管道通不通,343 条判据已经免费覆盖了。

按这条界,**目标闸的 live 判据是行为类的**:离线判据已经证明判断器拿得到 `read_file`、
越界会被驳回、判不出来时不假装成功;它们证明不了的是**这个具体模型拿到只读工具之后
Expand Down
12 changes: 7 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

[![tests](https://github.com/Jerry-TZ/Talos/actions/workflows/test.yml/badge.svg)](https://github.com/Jerry-TZ/Talos/actions/workflows/test.yml)

**一个你能完整读完的编程 agent。** 5311 行 Python,337 条离线判据,一份不糊弄人的安全说明。
**一个你能完整读完的编程 agent。** 5435 行 Python,343 条离线判据,一份不糊弄人的安全说明。

<img src="docs/demo.svg" alt="Talos 终端界面:动盘之前先弹确认框" width="100%">

Expand All @@ -19,8 +19,8 @@

| 文件 | 行数 | 职责 |
|---|---|---|
| `agent.py` | 4209 | 循环 + 工具 + 权限门 + 自学习 |
| `console_ui.py` | 214 | 终端界面(可整体替换) |
| `agent.py` | 4312 | 循环 + 工具 + 权限门 + 自学习 |
| `console_ui.py` | 235 | 终端界面(可整体替换) |
| `recall.py` | 582 | 联想记忆:扩散激活检索 |
| `session.py` | 306 | 会话持久化(想换 SQLite 只改这个) |

Expand All @@ -33,7 +33,7 @@
极简 agent 赛道很挤,有人用 Zig 做到 678KB 二进制。**Talos 不比谁小,它比谁都好读。**

- **能读完** — 四个文件,注释解释的是*为什么*,不是*是什么*。几乎每条防御旁边都写着它挡的那次真实翻车。
- **能验证** — 337 个测试,**离线、免 API key、几秒跑完**,CI 在 Linux/Windows × Python 3.10/3.13 上都跑。clone 下来立刻知道它没坏。
- **能验证** — 343 个测试,**离线、免 API key、几秒跑完**,CI 在 Linux/Windows × Python 3.10/3.13 上都跑。clone 下来立刻知道它没坏。
- **不吹牛** — [`SECURITY.md`](SECURITY.md) 明写 `create_tool` 就是进程内 RCE、正则黑名单只是减速带。**没有沙箱就是没有沙箱** —— 真要隔离,[三条现成方案](SECURITY.md#真要隔离怎么办)按代价从低到高列在那儿。
- **有考卷** — [`EXAM.md`](EXAM.md) / [`EXAM2.md`](EXAM2.md) 是两份可复现的能力测试,带标准答案和作弊检测(比如逐个核验 arXiv ID 真伪,防止编造引用)。记录的是"我怎么验证它真的有用",不是功能列表。
- **有实测** — [`FINDINGS.md`](FINDINGS.md) 记了二十二个真实任务量出来的东西:哪条提示词生效、哪条从头到尾没生效、六次翻车、检索改动的前后数字,以及**两个被数据否掉的自己的方案**。样本小,局限写在最前面。
Expand Down Expand Up @@ -159,6 +159,8 @@ tail -3 .talos/recall_trace.jsonl

**一次性模式** — `agent.py -p "任务"` 跑完即退,方便脚本和计时。

**流式(试验,默认关)** — `TALOS_STREAM=1` 让主循环边生成边收,转圈那一行实时显示「模型在写 write_file(src/app.py) · 12,000 字」,分得清慢和卡死。拼回来的回复跟整份返回的逐项相同(判据钉着),断流按原来的重试规矩整次重来。几家的流式形状只照文档和造出来的数据核对过,**还没真调用过** —— 跑出问题关掉,就回到原来那条路。

**可选自动化**(默认关闭,在 `talos.bat` 里开):

```bat
Expand Down Expand Up @@ -200,7 +202,7 @@ Ctrl+C 停下当前这轮 —— 做过的都留着,可以直接说

```bash
.venv\Scripts\python.exe -m pip install -r requirements-dev.txt
.venv\Scripts\python.exe -m pytest tests/ -q # 337 条,约 8 秒,不联网、不需要 key
.venv\Scripts\python.exe -m pytest tests/ -q # 343 条,约 8 秒,不联网、不需要 key
python agent.py --selfcheck # 免依赖的冒烟检查
```

Expand Down
105 changes: 104 additions & 1 deletion agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@
import re
import sys
import time
import types

# .env is loaded from the launch directory, which in a coding agent is very often somebody
# else's repository. Config is fine to pick up there; a command to execute is not — that turns
Expand Down Expand Up @@ -260,6 +261,13 @@ def _env_block() -> str:
# 后面也没法用。宁可这条稍紧,反正连不上会当场说清楚原因,不像超时那样一声不吭。
CONNECT_TIMEOUT = float(os.environ.get("TALOS_CONNECT_TIMEOUT", "5"))
SLOW_CALL = float(os.environ.get("TALOS_SLOW_CALL", "15")) # 超过这么久的调用才报耗时
# 流式:主循环那一次调用边生成边收,转圈那一行实时报「在写 write_file(x.py) · 12,000 字」。
# **默认关**:六家的流式细节各不一样(Gemini 不给 index、Kimi 把用量塞进 choices[0]),
# 写它的时候手上没有 key,只照文档和造出来的块核对过。开着跑稳了再改默认。
# 只影响看得见的那一处 —— 判断器、压缩摘要、复盘你看不到,流了也没用。
STREAM = os.environ.get("TALOS_STREAM", "").strip() in ("1", "true", "yes", "on")
# 不带 include_usage,流式默认**不回用量**:会话预算提醒和缓存命中率就静悄悄地全是 0
_STREAM_KW = {"stream": True, "stream_options": {"include_usage": True}}

ui = None # 界面 handle, set by repl(); kept out of module scope so --selfcheck is dep-free
_RUNTIME = {} # live client/model/state (+ subagent depth), set in agent_turn so tools like
Expand Down Expand Up @@ -2436,6 +2444,97 @@ def _retry_after(e):
except (TypeError, ValueError):
return None

def _ns(v):
"""dict → 能按属性读的对象,递归。Kimi 把流式的用量放在 `choices[0].usage`,SDK 不认识
这个位置,给的是原样的 dict —— 而 `_usage` 按属性读,读 dict 只会静悄悄地拿到 0。"""
if isinstance(v, dict):
return types.SimpleNamespace(**{k: _ns(x) for k, x in v.items()})
return v

# 参数是 JSON:Windows 路径在里面是 `src\\app.py`,抠的时候得认转义,抠完再解一遍
_PATH_ARG = re.compile(r'"path"\s*:\s*"((?:[^"\\]|\\.)*)"')

def _collect(stream):
"""把一次流式调用拼回**跟整份返回一模一样**的形状:`resp.choices[0].message` 带
content / tool_calls / reasoning_content,`resp.usage` 带用量。下游那几千行一行不用改 ——
流式换的只是「怎么拿到这份回复」,不是「回复是什么」。

工具调用是拼的重点,几家拆法不一样:
- OpenAI / DeepSeek:每片带 `index`,第一片给 id 和名字,后面只有一段段参数
- Gemini 3:**不给 `index`**,每个调用一整块,`thought_signature` 挂在这一块的
`extra_content` 上(丢了下一轮当场 400,见 `_tool_call_entry`)
- 有的每一片都重复同一个 id —— 那还是同一个调用
所以有 index 按 index 归位;没有就看 id:变了是新调用,没变或没给是接着上一个。
名字是**赋值**不是拼接(没有哪家把名字拆开发,而重复发名字的有)。

一边拼一边报进度(`ui.progress`),最该报的是写大文件:参数就是整份文件,要写几分钟。
流收完、断掉、被 Ctrl-C 打断,连接都关。"""
text, think, calls, at, usage = [], [], [], {}, None
n_text = n_think = 0
try:
for chunk in stream:
usage = getattr(chunk, "usage", None) or usage
for ch in (getattr(chunk, "choices", None) or ())[:1]:
usage = getattr(ch, "usage", None) or usage # Kimi 放在这儿
d = getattr(ch, "delta", None)
if d is None:
continue
said = None
r = _reasoning(d)
if r:
think.append(r)
n_think += len(r)
said = f"在思考 · {n_think:,} 字"
if getattr(d, "content", None):
text.append(d.content)
n_text += len(d.content)
said = f"在写回复 · {n_text:,} 字"
for tc in getattr(d, "tool_calls", None) or ():
i, cid = getattr(tc, "index", None), getattr(tc, "id", None)
f = getattr(tc, "function", None)
if i is not None and i in at:
c = calls[at[i]]
elif i is None and calls and not (cid and cid != calls[-1]["id"]):
c = calls[-1]
else:
c = {"id": cid, "name": "", "args": [], "n": 0, "path": None, "extra": {}}
calls.append(c)
if i is not None:
at[i] = len(calls) - 1
c["id"] = c["id"] or cid
if getattr(f, "name", None):
c["name"] = f.name
a = getattr(f, "arguments", None)
if a:
c["args"].append(a)
c["n"] += len(a)
# path 只在参数开头找:在后面的话要一遍遍 join 整份文件去搜,不值
if c["path"] is None and c["n"] - len(a) < 300:
m = _PATH_ARG.search("".join(c["args"])[:300])
try:
c["path"] = json.loads(f'"{m.group(1)}"') if m else None
except ValueError:
c["path"] = m.group(1)
c["extra"].update({k: v for k, v in
(getattr(tc, "model_extra", None) or {}).items()
if v is not None})
where = f"({c['path']})" if c["path"] else ""
said = f"在写 {c['name']}{where} · {c['n']:,} 字"
if said and ui is not None:
ui.progress(said)
finally:
close = getattr(stream, "close", None)
if callable(close):
close()
# 没给 id 的补一个:工具结果靠 `tool_call_id` 对回调用,None 发回去是 null,下一轮 400
tool_calls = [types.SimpleNamespace(
id=c["id"] or f"call_{k}", type="function", model_extra=c["extra"],
function=types.SimpleNamespace(name=c["name"], arguments="".join(c["args"])))
for k, c in enumerate(calls)] or None
msg = types.SimpleNamespace(content="".join(text) or None, tool_calls=tool_calls,
reasoning_content="".join(think) or None)
return types.SimpleNamespace(choices=[types.SimpleNamespace(message=msg)], usage=_ns(usage))

def _chat(client, **kwargs):
"""Call the model, retrying briefly on rate-limit / 'busy' / transient errors.
Free tiers (esp. glm-4.7-flash) get congested — '当前模型用户多' is just a busy signal."""
Expand All @@ -2444,6 +2543,10 @@ def _chat(client, **kwargs):
t0 = time.time()
try:
resp = client.chat.completions.create(**kwargs)
# 流式的错**在读到一半时**才出来,所以拼这一步必须在 try 里面:断流跟连不上
# 走同一套认法(掉线是 "connection",重试),重来那一趟从零拼,半截不带过来。
if kwargs.get("stream"):
resp = _collect(resp)
# 慢调用才报。快的不值一行,而推理模型上"这一次比平常久得多"正是你想知道的事。
took = time.time() - t0
if ui is not None and took >= SLOW_CALL:
Expand Down Expand Up @@ -3014,7 +3117,7 @@ def _seal_trace(steps: int, capped: bool) -> None:
resp = _chat(client, model=model,
messages=([{"role": "system", "content": system}]
+ _with_recall(messages, recalled, slot)),
tools=tool_specs())
tools=tool_specs(), **(_STREAM_KW if STREAM else {}))
state["tok"]["steps"] += 1
_in, _out, _cached = _usage(resp)
for _k, _v in zip(("in", "out", "cached"), (_in, _out, _cached)):
Expand Down
23 changes: 22 additions & 1 deletion console_ui.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,7 @@ def read_task(mode: str) -> str:
return console.input(f"[bold {C_YOU}]你[/] [dim]({mode})[/] › ").strip()

HEARTBEAT = 30 # 每这么多秒往下追加一行"还活着"
_active = None # 正在转的那个圈 —— `progress()` 往它上面写;同一时刻只有一个(子 agent 在圈外跑)

class _Thinking:
"""转圈 + 每 HEARTBEAT 秒**追加**一行「已等 Ns」。
Expand All @@ -68,8 +69,11 @@ def __init__(self, label: str):
self._stop = threading.Event()
self._status = None if console.legacy_windows else console.status(label, spinner="dots")
self._label = label
self._said = "" # 流式时模型写到哪了,`progress()` 填

def __enter__(self):
global _active
_active = self
if self._status is not None:
self._status.__enter__()
else:
Expand All @@ -89,9 +93,20 @@ def _beat(self) -> None:
waited += HEARTBEAT
# 用 :.0f —— waited 是累加出来的,HEARTBEAT 非整数时会攒出
# 0.15000000000000002 这种。生产里 HEARTBEAT=30 看不出来,测试里一眼就露。
console.print(f"[dim] … 已等 {waited:.0f}s,还在等模型回话(Ctrl-C 停这一轮)[/]")
# 流式时换成模型写到哪了 —— legacy_windows 不开转圈,这一行是那儿唯一的进度。
doing = f"模型{escape(self._said)}" if self._said else "还在等模型回话"
console.print(f"[dim] … 已等 {waited:.0f}s,{doing}(Ctrl-C 停这一轮)[/]")

def progress(self, said: str) -> None:
self._said = said
if self._status is not None:
# 转圈的**文字**原地换,不追加 —— 秒数那条说过原地重绘会刷屏,但那是
# legacy_windows 上;这里 `_status` 存在就说明不是那种控制台。
self._status.update(f"[{C_MODEL}]模型{escape(said)}[/]")

def __exit__(self, *exc):
global _active
_active = None
self._stop.set()
return self._status.__exit__(*exc) if self._status is not None else False

Expand All @@ -100,6 +115,12 @@ def thinking():
"""上下文管理器:模型思考时转个圈,并定期报"还活着"。"""
return _Thinking(f"[{C_MODEL}]模型思考中…[/]")

def progress(said: str) -> None:
"""流式时模型写到哪了(「在写 write_file(x.py) · 12,000 字」),显示在转圈上和心跳里。
转圈之外调它什么也不做。"""
if _active is not None:
_active.progress(said)

def took(seconds: float) -> None:
"""调用回来之后报一次耗时 —— 跟上面的心跳配套:心跳说"还活着",这条说"花了多久"。"""
console.print(f"[dim] ⏱ {seconds:.0f}s[/]")
Expand Down
3 changes: 3 additions & 0 deletions tests/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,9 @@ def _no_test_ever_writes_the_real_talos(tmp_path, monkeypatch):
# 就会往真实的 `.talos/sessions/` 里塞垃圾,而那批文件正是往事检索和燃尽表的语料。
# 上面那段说「枚举永远落后一步」,这就是下一步:**加功能等于给这片面新开一条路。**
monkeypatch.setattr(session, "SESS_DIR", os.path.join(d, "sessions"))
# 开着 `TALOS_STREAM=1` 的终端里跑 pytest,几百条用例的假客户端会收到 stream=True
# 而照旧回整份 —— 全套红一片,红的原因跟被测的东西毫无关系。要测流式的自己打开。
monkeypatch.setattr(agent, "STREAM", False)

@pytest.fixture(autouse=True)
def _keys_stay_put():
Expand Down
Loading
Loading