Best way to extract data from PDFs for a RAG pipeline #430
|
Hi everyone, However, the extraction part is nowhere to be seen... Do you have a best practices for extraction? I am using .NET so I leaned towards using pdfpig, document intelligence and perhaps content understanding if my PDF corpus contains lots of pictures. Any tips would be more than useful, cheers |
Replies: 2 comments
|
All right so I figured:
Above works great if all docs = mostly text + follow same layout. Otherwise:
Caveat: CU = ≤ 200 MB≤ 300 pages≤ 10 MB≤ 5 pages CU is generally cheaper than DI One more thing, DI is great when the layout/ format of your docs is static and when fields are alike across docs However, when they keep changing, it's a moving target CU is great at handling different formats, document types, different content.... I hope that helps |
|
Feel free to look at this repo where I implement both CU and DI https://github.com/Gabegi/agentic-retrieval-rag-production-azure-dotnet |
All right so I figured:
Above works great if all docs = mostly text + follow same layout.
Otherwise:
Caveat: CU = ≤ 200 MB≤ 300 pages≤ 10 MB≤ 5 pages
CU is generally cheaper than DI
One more thing, DI is great when the layout/ format of your docs is static and when fields are alike across docs
However, when they keep changing, it's a moving target
CU is great at handling different formats, document types, different content....
I hope that helps