CLI tool that builds training data for fine-tuning Spanish embedding models used in RAG
and agentic tool-calling retrieval. It takes source documents (PDF / Word), converts them to
Markdown, and generates (query, answer, hard_negative) triplets grounded in the specific
facts of each passage — so a query actually discriminates its passage from similar ones,
which is what a contrastive/triplet embedding objective needs.
The output CSV feeds the sibling repo Tesis-Embeddings (the trainer). This repo does not
do retrieval evaluation — that lives in Tesis-RAG / Tesis-Agent.
Each step is a CLI command reading/writing plain files on disk — no shared state.
flowchart LR
PDF[("PDF / DOCX<br/>source documents")]
MD[("Markdown *.md")]
CSV[("embeddings_qa.csv<br/>query · answer · hard_negative<br/>source_file")]
ENR[("embeddings_qa.csv<br/>+ hard_negative_mined")]
TE[("Tesis-Embeddings/<br/>datasets/*.csv")]
TR[("Tesis-RAG/<br/>dataset/ragval_dataset.csv")]
PDF -->|"pdf2md.py convert<br/>(docling)"| MD
MD -->|"synthetic.py create_embeddings<br/>(distilabel + Ollama/OpenAI)<br/>+ quality gate"| CSV
CSV -->|"synthetic.py mine_negatives (optional)<br/>(sentence-transformers)"| ENR
ENR -->|"copy into"| TE
CSV -->|"synthetic.py export_ragval"| TR
classDef data fill:#2b3a55,stroke:#12213a,color:#fff
class PDF,MD,CSV,ENR,TE,TR data
create_embeddingsruns against any OpenAI-compatible endpoint: setbase_urlin its config to an Ollama server (the default, fully local) or leave it empty and exportOPENAI_API_KEYto use OpenAI.mine_negativesdownloads asentence-transformersbaseline model on first run.pdf2md convertis fully local (docling).
- Python 3.12 or higher
- Required libraries (see
requirements.txt):jsonargparse,tqdm,pandasdocling— PDF/DOCX → Markdown conversiondistilabel[openai]— grounded triplet generation (create_embeddings)sentence-transformers,datasets— corpus hard-negative mining (mine_negatives)
OPENAI_API_KEYexported in the environment forcreate_embeddings
git clone https://iieg-app.jalisco.gob.mx/iieg-ia/llm-synthetic-data.git
cd llm-synthetic-data
pip install -r requirements.txt💡 Every command reads its parameters from a YAML in
configs/. Edit the config rather than passing flags.
Converts every PDF and Word (.docx) file under the input directory to Markdown (via
docling), which is cleaner input for the LLM than raw PDF/DOCX.
python pdf2md.py convert --config configs/pdf2md.yamlGenerates (query, answer, hard_negative) triplets — one query per paragraph, explicitly
grounded in that paragraph's facts/entities, plus a hard negative authored alongside it.
distilabel's GenerateSentencePair owns the prompt and the parsing. Writes
embeddings_qa.csv.
⚠️ Do not turnuse_default_structured_outputback on for an Ollama-served model: its tool-calling mode collapses on real anchors — measured 0/20 against 19/20 with the task's own parser, and 3.5× slower.
Generated pairs go through an automated quality gate before being written (the criteria of
protocol section 10.3): unsubstituted {placeholders}, meta-instructions to the generator,
queries with no lexical anchoring to their passage, queries anchored only in generic vocabulary
with no figure or proper noun, and exact duplicates. Rejects land in rejected_qa.csv and the
rejection rate is logged. The gate exists because the first evaluation set had to be discarded:
all 17 models scored at chance level, and re-running the gate over that dataset rejects 100 % of it.
# Local generation against Ollama (default config)
python synthetic.py create_embeddings --config configs/create_embeddings.yaml
# ... or against OpenAI: clear base_url in the config first
export OPENAI_API_KEY=sk-...
python synthetic.py create_embeddings --config configs/create_embeddings.yaml
⚠️ Reasoning models served by Ollama (e.g.qwen3.6) return an empty answer unless thinking is disabled — the whole token budget goes to thereasoningfield. Keepdisable_thinking: truefor them, andmax_new_tokenswell above distilabel's default of 128.
Adds a hard_negative_mined column — a real confusable passage from the corpus, found by
embedding similarity (usually a stronger negative than one the LLM imagines in isolation).
Tesis-Embeddings' trainer prefers this column when present.
python synthetic.py mine_negatives --config configs/mine_negatives.yamlCopy the resulting CSV into Tesis-Embeddings/datasets/ and train there.
Derives Tesis-RAG's schema (id, pregunta, chunk_id, chunk_content, documento) from the same
generated CSV, so one generation feeds both the trainer and the retrieval evaluator. chunk_id is
the same SHA-256 prefix Tesis-RAG computes, and documento carries the source file — which is
what lets Tesis-Embeddings split by document (--group-col source_file) instead of by row.
python synthetic.py export_ragval --config configs/export_ragval.yamlslurm/ holds one script per step. The whole pipeline (generation → training → both benchmarks)
is submitted as a single dependency chain from Tesis-Embeddings/slurm/submit_all.sh. Generation
needs a conda env with distilabel (datagen) and no GPU allocation: it talks over HTTP to the
ollama serve daemon already running on the node.
Contributions are welcome! Please feel free to open an issue for any suggestions or improvements.
This project is licensed under the MIT License - see the LICENSE file for details.
- docling: parses PDF, DOCX, XLSX, HTML and more, exporting to Markdown/HTML/JSON.
- distilabel: synthetic data generation framework; its
GenerateSentencePairtask produces the grounded triplets with enforced structured output. - sentence-transformers: embedding models and
mine_hard_negativesfor corpus-grounded negative mining.