Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Alexi Feng

I study how AI systems fail—and how evaluations can capture those failures more reliably.

My work focuses on AI evaluation, benchmark design, agent experiments, and reproducible failure analysis. I previously worked on Seed model evaluation at ByteDance and now continue this work as an independent researcher and builder.

Research focus

  • LLM and agent evaluation
  • Benchmark design and evaluation health
  • Failure analysis and reproducible experiments
  • Evaluation tooling and research workflows

Selected research

Writing

I publish research notes in English and Chinese.

Open to collaboration on AI evaluation, benchmark research, and agent reliability.

About

AI evaluation research, benchmark design, agent experiments, and reproducible failure analysis.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors