PDF Data Extraction at Scale: When to Trust an LLM

LLMs read chaotic documents like no one else, but when money or legal liability is on the line, the question is not "can it?" but "when should I trust it, and how do I catch it when it is wrong?". In this hands-on workshop we build, step by step, a real extraction pipeline for legal PDFs with variable formats, orchestrated with Airflow, using a hybrid extractor (deterministic + LLM) and deterministic guardrails in Python. You will leave with a framework for deciding which tool to use and a reliable pattern for production environments.

Want to know more?

Join PyCon Colombia newsletter and get a complete overview of our events, speakers and community participation.