At a glance
Check for a CSV export or API first. If PDFs are the only option, identify whether they contain embedded text or scanned images. Local tools can extract either, but the results still need checks before they update your records.
Common import errors
Scroll the table to see all columns.
| Example problem | Check | Expected outcome |
|---|---|---|
| Account 00127 becomes 127 | Preserve identifiers as text; match to the source system | Keep the leading zeros |
| An amount shifts into the next row | Match identifiers and fields together; inspect source position | Hold the affected record for review |
| The same report is uploaded twice | Check file fingerprint and business record keys | Show the previous import; avoid duplicate records |
| A revised report changes an amount | Compare against the existing record and revision | Show the difference before replacing accepted data |
Choose the extraction method
Try copying a line from the PDF into a text editor. Check it against the page. Selectable text may come from an earlier OCR pass, so it can still contain mistakes.
pypdf reads embedded text but does not perform OCR. Tables can be difficult because the PDF may store text positions without row and column structure. pdfplumber provides table extraction and tools for inspecting those positions.
Scans need OCR. OCRmyPDF adds a searchable text layer. It does not determine which account an amount belongs to or whether the amount is correct.
Test several reports, including a long one, a corrected one, and a record split across pages. One clean sample won’t show the variations your staff handle.
Define the fields before importing
Keep account numbers as text so 00127 stays 00127. Decide whether a blank amount means “not supplied” rather than zero. Confirm the report’s date convention instead of guessing what 03/04/2026 means.
Store a reference to the source file and page with each record. Include the parser version so you can identify affected imports if a problem is found later.
Check records and totals
An amount can be read correctly but attached to the wrong account. Check identifiers and values together. Match record counts and totals to the report where available, but remember that swapped amounts can still produce the right total.
Allow for valid credits, adjustments, and rounding. Send uncertain records for review with a specific reason, such as “two account matches” or “total differs by $18.50.” Show the source beside the proposed value.
Handle duplicates and corrections
A file fingerprint detects an identical upload. A new PDF may contain the same records, so also check the source system’s record identifiers. A filename or row number is not a reliable record key.
Treat a correction separately. Show the changed fields before replacing accepted data, and retain the earlier version when the business needs a history.
Interrupt an import during testing. It should report what was saved and resume without creating those records again.
Measure accuracy and review time
Have a person check a reference set field by field. Reserve some files for testing rather than using every example to develop the parser.
Suppose 900 of 1,000 records are accepted automatically, but checking finds nine errors in that group. Automatic acceptance is 90%; the error rate among accepted records is 1%. Report both.
Measure the review work too. At five minutes each, 100 flagged records take eight hours and 20 minutes. Use that figure when estimating savings.
Check where files go
Local processing still needs access controls, backups, and retention rules. Check temporary files and error logs as well as the originals. OCRmyPDF’s security guidance also makes clear that it is not malware protection.
For Eagle Medical Billing, we built local parsing, record checks, and a follow-up queue because the billing system offered PDFs without an export or API. Include that review work in the project budget, not just the extraction.