Description
When running FLEX evaluation with test_flex.py on math tasks, the process can hang indefinitely with one Python thread stuck at ~100% CPU.
This seems to happen when the underlying smolagents.CodeAgent generates Python code that:
- enters a very heavy computation path, or
- effectively times out / gets stuck in a long-running loop.
In this case, the whole batch can stop making progress because asyncio.gather(...) waits for all samples in the batch to finish.
Environment
- FLEX repo: GenSI-THUAIR/FLEX
- Task: math / AIME
- Command example:
python test_flex.py \
--task_type math \
--actor deepseek-chat \
--data_path ./data/AIME/ \
--split test \
--batch_size 4 \
--no-retrieve \
--no-telemetry \
--results_dir results/math_baseline_full_rerun
Observed behavior
The process appears "stuck"
One Python thread stays at ~100% CPU
No new result files are produced for a long time
The current batch does not finish, so later samples never start
Suspected cause
The issue appears related to the local smolagents code executor used by CodeAgent.
FLEX creates a CodeAgent in actor.py, and each sample calls:
result = await asyncio.to_thread(self.agent.run, prompt)
If the generated code becomes pathological, the local executor timeout does not reliably recover the worker, and the sample can effectively block the entire batch.
This is especially problematic in:
actor.py
test_flex.py
Reproduction hint
AIME-style recurrence / large-integer problems seem especially likely to trigger this behavior, because the agent may generate exact rational recurrences that blow up in integer size.
Expected behavior
A single bad sample should not block the whole batch indefinitely
Per-sample execution should fail fast after timeout
Evaluation should continue to later samples
Possible fixes
Add a hard per-sample timeout at the process level
Avoid waiting for an entire batch with asyncio.gather(...) if one sample is stuck
Mark timed-out samples as failed and continue
Consider isolating code execution more robustly than the current local threaded executor
Description
When running FLEX evaluation with
test_flex.pyon math tasks, the process can hang indefinitely with one Python thread stuck at ~100% CPU.This seems to happen when the underlying
smolagents.CodeAgentgenerates Python code that:In this case, the whole batch can stop making progress because
asyncio.gather(...)waits for all samples in the batch to finish.Environment
python test_flex.py \ --task_type math \ --actor deepseek-chat \ --data_path ./data/AIME/ \ --split test \ --batch_size 4 \ --no-retrieve \ --no-telemetry \ --results_dir results/math_baseline_full_rerunObserved behavior
The process appears "stuck"
One Python thread stays at ~100% CPU
No new result files are produced for a long time
The current batch does not finish, so later samples never start
Suspected cause
The issue appears related to the local smolagents code executor used by CodeAgent.
FLEX creates a CodeAgent in actor.py, and each sample calls:
result = await asyncio.to_thread(self.agent.run, prompt)
If the generated code becomes pathological, the local executor timeout does not reliably recover the worker, and the sample can effectively block the entire batch.
This is especially problematic in:
actor.py
test_flex.py
Reproduction hint
AIME-style recurrence / large-integer problems seem especially likely to trigger this behavior, because the agent may generate exact rational recurrences that blow up in integer size.
Expected behavior
A single bad sample should not block the whole batch indefinitely
Per-sample execution should fail fast after timeout
Evaluation should continue to later samples
Possible fixes
Add a hard per-sample timeout at the process level
Avoid waiting for an entire batch with asyncio.gather(...) if one sample is stuck
Mark timed-out samples as failed and continue
Consider isolating code execution more robustly than the current local threaded executor