Reading an engineering drawing with optical character recognition (OCR) is not the same as extracting ordinary text from a PDF. A dimension may combine a nominal value, tolerance, symbol, quantity, thread class, leader, geometric dimensioning and tolerancing (GD&T) feature-control frame, or note whose meaning depends on its location. Useful extraction therefore needs character recognition plus spatial grouping, drawing-specific interpretation, and reviewer verification.
First ask: how is the page stored?
Contents
A PDF page can contain different kinds of content, even within the same drawing:
- Native text objects: characters exist as text and may be selectable. Their order and relationship to nearby geometry can still be wrong.
- Vector outlines: text looks crisp when zoomed but is stored as drawn shapes rather than searchable characters.
- Raster scans: the page is pixels. OCR must infer characters from the image.
- Mixed pages: text, vector geometry, and embedded scans coexist.
Try selecting and copying a callout, then zoom in. Crisp appearance does not prove the text is extractable: outlined letters also stay crisp. A PDF inspection library can distinguish text objects from image-heavy pages, but automated classification can be uncertain.

Why ordinary OCR has trouble
Invoices and forms often follow a reading order: one text line after another. Engineering drawings do not. Their text is sparse but spatially meaningful; values may sit above, below, or beside dimension lines and leaders. Notes can apply globally, symbols can modify values, and the same feature can appear in multiple views.
| Drawing content | Common failure | Why the relationship matters |
|---|---|---|
| Decimal values | 8.0 read as 80; decimal point disappears | A small mark changes the allowed size |
| Diameter/radius | Ø confused with 0 or O; R omitted | Symbol changes the kind of requirement |
| Stacked tolerance | Upper/lower values merged or swapped | Limits can become asymmetric |
| Limit dimensions | Two limits read as separate characteristics | Both numbers define one acceptable range |
| Rotated dimensions | Text orientation or reading order is wrong | The callout may attach to the wrong feature |
| Thread notation | Pitch, class, depth, or quantity is lost | The leftover string can be incomplete |
| Surface-finish symbol | Symbol omitted or mistaken for linework | Acceptance intent may disappear |
| GD&T frame | Cells or datum order are read incorrectly | The tolerance meaning depends on the complete frame |
Ground truth: correct text is not enough
Take this synthetic callout: 4X Ø8.0 ±0.1 THRU. A text recognizer might return the same characters correctly but place them next to the wrong leader. It might find “4X” and “Ø8.0 ±0.1” separately but fail to group them. It might group the size correctly and miss a nearby position control. Each case has different consequences, even though a character-level score may look good.
| Ground-truth item | Possible extraction | What failed? |
|---|---|---|
| 4X Ø8.0 ±0.1 THRU, upper-right pattern | “4X 08.0 ±0.1 THRU” at the correct source | Symbol recognition |
| 4X Ø8.0 ±0.1 THRU, upper-right pattern | “4X” and “Ø8.0 ±0.1” as two candidates | Grouping |
| 4X Ø8.0 ±0.1 THRU, upper-right pattern | Perfect text attached to lower-left hole | Location association |
| Position ⌀0.30 | A | B | C | “0.30” without symbol or datums | Semantic content lost |
Keep raw text, source location, grouped interpretation, and reviewer decision distinct. A candidate with correct text but wrong grouping is not a correct characteristic record.
Vector extraction, OCR, layout parsing, and vision
| Method | Strength | Limit | Useful role |
|---|---|---|---|
| Native PDF text extraction | Can preserve encoded characters | Reading order and grouping may be wrong; outlines are not text | Digital PDFs with text objects |
| Raster OCR | Can read scanned or rendered regions | Can misread symbols, small text, skew, and decimals | Scanned, flattened, or outlined text |
| Geometry/layout analysis | Uses page coordinates and relationships | Requires drawing-specific logic | Grouping callouts and locating features |
| Vision models | Can use broader visual context | May be variable, costly, or unsuitable under data rules | Assisted interpretation with review |
| Hybrid workflow | Uses methods where they fit | More components; still needs review | Production process with evidence checks |
PyMuPDF describes OCR as a slower path for image-based or text-poor pages. Its current Page API also says that, since version 1.27.2, partial OCR processes page areas outside legible text, including vector graphics. This can recognize lettering after it is rendered to pixels; it does not make OCR understand line-art geometry, leaders, or feature relationships. Use native text and geometry extraction where available, then verify each grouped callout against the drawing.
Confidence is not correctness
A confidence score estimates something about the detection system; it is not an acceptance decision. OCR can confidently read 8.00 instead of 6.00. It can read the characters correctly but attach them to the wrong feature. It can omit the tolerance or fail to capture a global note. Reviewers need the source crop, surrounding geometry, and full grouped requirement—not just a score or green status.
How to assess OCR and extraction claims
Before trusting a published accuracy number, ask what it measured:
- Were pages vector, scanned, or mixed, and how clean were they?
- Were decimals, tolerances, symbols, threads, notes, and GD&T represented?
- Was correctness scored per character, token, callout, or complete inspection characteristic?
- Did the test require correct grouping and source location?
- How were missed items and false positives counted?
- Was performance measured before or after reviewer corrections?
- Was the test set independent of the software’s tuning examples?
A study of DigiEDraw, a research prototype for extracting dimension requirements, discusses the difficulty of maintaining spatial relationships among drawing text and geometry. Its dataset and evaluation belong to that study; do not transfer its results to another product or to every engineering drawing. Read the DigiEDraw paper and the earlier work on dimensioning-text segmentation and recognition for the research context.
A practical hybrid pipeline
- Classify the page and regions as text objects, outlines, raster, or mixed.
- Extract native text and coordinates where available.
- Render or rescan difficult regions at sufficient resolution and orientation.
- Group text, symbols, tolerances, leaders, and geometry into candidate requirements.
- Retain source location and alternate readings where supported.
- Send uncertain or high-impact candidates to a reviewer.
- Preserve raw source and reviewer-corrected values separately.
- Validate the drawing-to-report mapping before exporting.
How BalloonForge supports review
BalloonForge combines native PDF text, raster and rotated OCR, grouping logic, GD&T-frame geometry, and localized rescanning. Its correction view shows the source crop alongside raw and normalized readings, and lets the reviewer correct, include, reject, or note a candidate. GD&T frames are detected as regions, but current product documentation says not all individual cells are semantically decoded. Review the complete feature-control frame against the drawing.

BalloonForge processes drawing and project data locally on the Windows device according to its current product documentation. The app also has privacy-filtered exception reporting. Local drawing processing should not be confused with “no telemetry”; review the current BalloonForge privacy policy and your organization’s rules before processing restricted files.
Frequently asked questions
Can OCR read text in a vector PDF?
It can sometimes recognize lettering after the page or region is rendered to pixels. Native text extraction may preserve encoded characters directly, while OCR does not by itself interpret vector geometry, leaders, or feature relationships. Identify how the content is stored and verify each callout and its location.
Can OCR read GD&T?
It may recognize some characters, but the frame’s compartments, symbols, modifiers, datum references, and relationship to the feature need verification.
Can OCR extract tolerances?
It can detect tolerance text, but grouping and sign/order errors are possible. Check the source callout and resolve any governing general tolerance before using the value.
Can OCR find every inspection characteristic?
No completeness claim should be assumed. Use a known review procedure and inspect notes, patterns, repeated views, and low-confidence regions.
Further reading
Free engineering resource
Drawing Inspection Starter Kit
An editable Excel report template, a printable drawing-review checklist, and a guide to connecting drawing requirements with inspection results.
ZIP · Excel template and printable HTML guides
Enter your email and submit the form to unlock your download.




