We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm.
# Clone the repository
git clone https://github.com/GenSI-THUAIR/AMix-1.git
cd AMix-1
# Create and activate a Python 3.10 environment (recommended: conda)
conda create -n amix python=3.10 -y
conda activate amix
# Alternatively, using venv
# python3.10 -m venv venv
# source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtDownload config.yaml and model checkpoint through Huggingface:
# login through your HuggingFace key
huggingface-cli login
# run code in src
cd src
# download the whole AMix-1-1.7B directory
huggingface-cli download GenSI/AMix-1-1.7B --local-dir ./AMix-1-1.7B --local-dir-use-symlinks False
# adjust directory structure
mv AMix-1-1.7B/.hydra AMix-1-1.7B/ckpt ./
# back to AMix-1
cd ..Before inference, the directory structure should be:
AMix-1/
├── imgs
├── src
│ ├── .hydra
│ │ └── config.yaml
│ ├── ckpt
│ │ └── AMix-1-1.7b.ckpt
│ ├── model.py
│ ├── inference.py
│ └── ...
│
├── README.md
├── inference.sh
└── ...
Parameters for inference.sh
| Parameter | Description | Example |
|---|---|---|
--input_seq |
Input MSA, a single sequence or multiple sequences with equal length | "ASAAA" or "AAASA,SAASA,ASASA" |
--output_dir |
Sequences generated per round | "./output" |
--num_seq |
Number of sequences to generate | 10 |
--time |
Noise factor for generation (0.0 to 1.0), 1.0 means no noise | 0.8 |
--ckpt_path |
Checkpoint path for the model | "./ckpt/AMix-1-1.7b.ckpt" |
--seed |
Random seed for reproducibility | 42 |
bash inference.sh --input_seq "AAASASA" --output_dir "./output" --num_seq 10 --time 0.8 --ckpt_path "./ckpt/AMix-1-1.7b.ckpt" --seed 42
EvoAMix-1 is our test-time scaling algorithm that iteratively evolves protein sequences through generation, evaluation, and selection cycles. This approach enables continuous improvement of protein designs by leveraging multiple evaluation metrics and sophisticated filtering strategies.
We provide a convenient bash script run_tts.sh to launch the EvoAMix-1 pipeline with proper parameter configuration:
./run_tts.sh --rounds 10 --num-seqs 50 --top-k 5 --eval-task-weights "pLDDT:1.0"Before running, you must configure these essential paths in run_tts.sh:
# TODO: Set your model checkpoint file path
CKPT_FILE="./ckpt/AMix-1-1.7b.ckpt"
# TODO: Set your initial data directory path
INIT_DATA="/path/to/your/data/CASP14_orphan"
# TODO: Set your experiment output base directory
EXP_BASE_DIR="/path/to/your/experiments"| Parameter | Description | Example |
|---|---|---|
--rounds |
Number of evolution iterations | 10 |
--num-seqs |
Sequences generated per round | 50 |
--top-k |
Top sequences selected for next round | 5 |
--eval-task-weights |
Evaluation metrics and weights | "pLDDT:1.0,rosetta_energy:-1.0" |
--eval-filter |
Filter criteria for sequences | "pLDDT>80,progen_nll<10" |
--infer-step |
Inference steps for generation | 50 |
--esm-fold-gpus |
Number of GPUs for ESM-Fold | 1 |
The evaluation weights use the format "metric1:weight1,metric2:weight2":
- Positive weights: Higher scores are better (e.g.,
pLDDT:1.0) - Negative weights: Lower scores are better (e.g.,
rosetta_energy:-1.0)
Filters use the format "metric1>value1,metric2<value2" and support operators: >, <, >=, <=, ==, !=
You can easily extend EvoAMix-1 with custom evaluation metrics by implementing new verifiers in src/evaluation/metrics.py. Each verifier should follow this pattern:
def evaluate_your_metric(eval_file, args=None):
"""
Evaluate sequences using your custom metric
Args:
eval_file: Path to FASTA file containing sequences
args: Optional command line arguments
Returns:
Dictionary mapping sequence IDs to scores
"""
logger.info("=== Starting your metric evaluation ===")
abs_eval_file = os.path.abspath(eval_file)
sequences = parse_fasta(abs_eval_file)
metric_scores = {}
for seq_record in sequences:
seq_id = seq_record.id
sequence = str(seq_record.seq)
# Implement your evaluation logic here
score = your_evaluation_function(sequence)
# Store score in annotations
if 'annotations' not in seq_record.__dict__:
seq_record.annotations = {}
seq_record['annotations']['your_metric'] = score
metric_scores[seq_id] = score
# Write updated sequences back to FASTA
write_fasta(sequences, abs_eval_file)
logger.info("=== Your metric evaluation completed ===")
return metric_scores- Function signature: Must accept
eval_fileand optionalargsparameters - Score annotation: Store scores in FASTA description as
metric_name=value - Sequence annotations: Store scores in
seq_record['annotations'][metric_name] - File update: Write updated sequences back to the original FASTA file
- Registration: Add your function to the
METRICSdictionary:
METRICS = {
# ... existing metrics
"your_metric": evaluate_your_metric,
}We provide several built-in evaluation metrics:
pLDDT: Protein local structure quality (via ESM-Fold)pTM: Protein template modeling score (via ESM-Fold)TM_score: Structural similarity using TM-alignrosetta_energy: Rosetta energy scoringnovelty: Sequence novelty compared to originaldiversity: Intra-sample sequence diversityrepeatness: Amino acid repeat analysisidentical: Sequence identity to original
Important: Different verifiers may have conflicting dependencies or specific environment requirements. To avoid conflicts, we strongly recommend:
Create separate conda/virtual environments for verifiers with complex dependencies:
# Example: Create environment for a deep learning verifier
conda create -n verifier_env python=3.10
conda activate verifier_env
pip install your_verifier_requirements
# Example: Create environment for structure-based verifiers
conda create -n structure_env python=3.9
conda activate structure_env
pip install pymol biopythonImplement verifiers that require special environments using subprocess calls:
def evaluate_external_verifier(eval_file, args=None):
"""Example of calling external verifier in isolated environment"""
logger.info("=== Starting external verifier evaluation ===")
# Call external script in specific environment
cmd = [
"conda", "run", "-n", "verifier_env", # Or your python path
"python", "/path/to/external_verifier.py",
"--input", eval_file,
"--output", eval_file # Update in-place
]
try:
result = subprocess.run(cmd, check=True, capture_output=True, text=True)
logger.info("External verifier completed successfully")
except subprocess.CalledProcessError as e:
logger.error(f"External verifier failed: {e}")
raise
# Parse results from updated FASTA file
sequences = parse_fasta(eval_file)
scores = {}
for seq in sequences:
# Extract score from description or annotations
scores[seq.id] = extract_score_from_seq(seq)
return scoresEvoAMix-1 generates comprehensive outputs in your experiment directory:
exp_dir/
├── tts.log # Main execution log
├── error.log # Error and warning log
├── data.json # Complete evolution history
├── round_1/
│ ├── filtered_sequences_round_1.json
│ └── tmp_round_1_all_samples.fasta
├── round_2/
│ └── ...
└── plots/ # Evolution curve visualizations
├── sample_0_metrics_plot.png
├── all_samples_weighted_score_plot.png
└── average_metrics_plot.png
The evolution process tracks all metrics across rounds, enabling detailed analysis of sequence improvement over time.
AMix-1 would like to express heartfelt thanks to the following projects and contributors.
This work builds upon and adapts ideas and code from:
- microsoft/evodiff, which provided the preprocessed UniRef50 dataset, tools for sequence sampling evaluation, and the data processing pipeline.
- bytedance/dplm, whose two samplers were utilized in inference.py.
- facebookresearch/esm, which inspired the design of the model architecture.
We deeply appreciate the efforts of the creators of these repositories, whose work has been instrumental in shaping AMix-1.
@article{lv2025amix1,
title={AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model},
author={Changze Lv*, Jiang Zhou*, Siyu Long*, Lihao Wang, Jiangtao Feng, Dongyu Xue, Yu Pei, Hao Wang, Zherui Zhang, Yuchen Cai, Zhiqiang Gao, Ziyuan Ma, Jiakai Hu, Chaochen Gao, Jingjing Gong, Yuxuan Song, Shuyi Zhang, Xiaoqing Zheng, Deyi Xiong, Lei Bai, Wanli Ouyang, Ya-Qin Zhang, Wei-Ying Ma, Bowen Zhou, Hao Zhou},
journal={arXiv preprint arXiv:2507.08920},
year={2025}
}
