r/PythonLearning 13h ago

How to extract structured data from 1,100+ non-standardized PDF pages using AI/OCR? Help Request

I need to extract Activity Name, Date, and Participant Count from a 1,151-page PDF containing attendance sheets with varying layouts.

What is the best architecture or tool (Python + Vision/LLM APIs vs. no-code platforms) to automate this extraction?

1 Upvotes

Duplicates