Skip to content
Discussion options

You must be logged in to vote

All right so I figured:

  • Cheap text/ metadata extraction = SDK (python/ .NET) e.g. PDFPig
  • Document Intelligence (DI) Layout option = Text + Structure + Metadata

Above works great if all docs = mostly text + follow same layout.

Otherwise:

  • Content Understanding (CU)

Caveat: CU = ≤ 200 MB≤ 300 pages≤ 10 MB≤ 5 pages

CU is generally cheaper than DI

One more thing, DI is great when the layout/ format of your docs is static and when fields are alike across docs

However, when they keep changing, it's a moving target

CU is great at handling different formats, document types, different content....

I hope that helps

Replies: 2 comments

Comment options

You must be logged in to vote
0 replies
Answer selected by Gabegi
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
1 participant