A table that looks perfectly clear on screen can arrive in a CAT tool as a jumble of labels, values, and footnotes. A two-column page can be extracted across both columns instead of down the first and then down the second. The page still looks orderly; the text handed to the translator does not.
The cause usually isn’t a translator pressing the wrong button. A PDF is designed to preserve a page’s appearance, not to act like a structured document ready for editing. A CAT tool needs extractable text and segments it can translate. Between the visible page and those segments, conversion has to infer what belongs together and what comes next.
That inference can fail before translation starts. If a heading gets attached to the wrong paragraph or a value gets separated from its table label, even a fluent translation can preserve the wrong relationship. Good PDF work starts with checking what the file contains and what the conversion actually recovered.
Why a PDF page isn’t the same as a structured document¶
A PDF page is a visual arrangement of text and other content. A table shown on that page isn’t necessarily stored as a table with rows, columns, and cell relationships. Likewise, text placed in two columns isn’t automatically marked with a logical instruction to read down one column and then the other.
That difference matters because CAT tools don’t necessarily translate the page as a person sees it. A tool may first convert the PDF into an editable intermediate format, such as DOCX, and then process that file. The intermediate document, rather than the original page image or layout, becomes the material the translator works on.
The visible positions of words can offer clues, but positions aren’t the same as structure. A line near the top of a page might be a heading, a running header, a label in a table, or a callout. A value aligned underneath another value might belong in the same column, or it might be the start of another visual region. Conversion software has to work out which interpretation makes sense.
Tagged PDFs can carry more useful information. Tags and a document structure can express the intended order of content and identify elements such as headings or tables. The W3C guidance on PDF reading order explains that correct reading order depends on the document’s structure, not only on where text appears on a page. The PDF Association’s guidance on accessible PDFs likewise describes logical order as something determined by a tag tree.
A tag tree doesn’t make every complex PDF easy to translate. Tags can be missing or poorly arranged, and a visual layout can be hard to represent as a simple sequence. Closely spaced columns and irregular text arrangements can confuse automatic tagging. Adobe’s accessible PDF guidance notes that complex structures and reading order may need manual verification or repair.
The practical distinction is simple: a page can look right and still lack the machine-readable relationships a conversion needs. A reviewer should treat the PDF’s appearance and its extracted structure as two different things to inspect.
How a CAT workflow can turn a layout into scrambled text¶
A PDF translation workflow often includes several transformations: extracting content, converting it into a format the CAT tool can process, translating the segments, and preparing a deliverable. Each step can preserve some parts of the page while changing or losing others.
The most important trade-off is between recovering text flow and preserving page appearance. A conversion that prioritizes continuous text may change the original formatting. A conversion that tries to reproduce the page more closely may preserve text positions but fail to recover some content or relationships. No single conversion choice can guarantee that the translated document will retain both the original appearance and a clean, logical text flow.
That trade-off helps explain why a translator might see one of several problems:
- Columns interleave. Text from the left and right columns appears in an alternating sequence, even though the reader would finish one column before moving to the next.
- Table cells lose their associations. A row label may be separated from the value it describes, or values may be extracted without their column headings.
- Text disappears or repeats. A conversion may miss content or interpret overlapping text as separate items.
- Line breaks create misleading segments. A sentence split across visual lines may turn into several fragments, or the conversion may join unrelated text.
- Headers, footers, and footnotes interrupt the body. Repeating page elements can appear between paragraphs or in the middle of a table.
- The translated file doesn’t look like the source. The CAT tool may return an editable document or extracted text instead of a reconstructed PDF.
A particularly risky choice is importing PDF content as plain text when the task requires a formatted deliverable. Plain-text extraction discards formatting, so it can’t provide the layout cues needed for a document intended to return as a structured, readable file. Plain text can still suit other tasks, such as working with text for corpus or alignment purposes, but that is a different goal from preserving the document’s layout.
The output format needs a separate conversation. A translator or project manager should not assume that importing a PDF means the CAT tool will export a translated PDF with the same appearance. A workflow may produce an editable file or plain text, and page reconstruction may need to happen outside the CAT tool.
Teams can avoid a painful handoff by agreeing on the intended deliverable before processing begins. Is the client asking for translated text, an editable DOCX, or a layout-matched PDF? Does the work include reconstruction and final visual review, or only translation? The answers change the conversion choices and the amount of checking required.
For practical evidence of the repair work involved, a ProZ discussion about PDF formatting problems describes a workflow where the converted file’s tables and layout were corrected before translation, then checked and repaired again afterward. The post is an individual practitioner’s experience, not a universal product specification. Its useful lesson is that file conversion, translation, and layout review are separate tasks.
Tables and multi-column pages need different checks¶
A table is not just text arranged into a grid. The reader needs to know which header applies to each column, which label belongs to each value, and how a row continues when a table crosses a page break. A conversion that keeps the words but loses those relationships can leave the translator with text that is technically present but hard to interpret safely.
For a simple table, compare the extracted content with the visible rows and columns. Check that each cell stays with the correct row, that values remain under the right headers, and that units or notes haven’t drifted away from the numbers they qualify. A table with merged cells, repeated headers, or nested information needs closer review because its relationships can’t always be inferred from spacing alone.
Long tables deserve a continuity check. When a table continues on a new page, confirm that the continuation still belongs to the same table, that repeated headers aren’t mistaken for new data, and that every value remains associated with the correct heading. The Tagged PDF Best Practice Guide highlights the importance of verifying cell-to-header relationships and continuity across page breaks in complex tables.
Consider a form that places a field label in one column and its value in the next. If extraction places the labels in one block and the values in another, the translator might be able to infer the intended pairs from the page, but the CAT segments no longer show those pairs reliably. A reviewer should resolve the structure before translation rather than hoping that context will repair it later.
Multi-column pages pose a different ordering problem. The extraction needs to reproduce the intended reading sequence, not simply collect text by its position across the page. A poorly tagged page can be interpreted across columns incorrectly, and complex layouts with sidebars, footnotes, graphics, or form fields may not be placed in the correct order automatically, according to the W3C PDF reading-order technique.
To check a two-column page, read the extracted text against the page as a person would. Confirm that the first column proceeds from top to bottom, then check where the second begins. Look for an interleaved sequence in which a heading from one column is followed by a paragraph from the other. Pay particular attention to sidebars and footnotes: those elements can be visually close to body text but belong elsewhere in the reading order.
A practical preflight checklist can make the review consistent:
- Compare the extracted sequence with the visible page, including each column’s start and end.
- Match table labels, values, units, and notes to the correct headers and rows.
- Check repeated headers and tables that continue across page breaks.
- Identify headings, footnotes, sidebars, and running page elements that interrupt the text.
- Mark any content that needs repair before translation, then recheck the repaired version.
The checklist doesn’t require a particular CAT tool. It requires a clear comparison between the original page and the file the translator will actually process. The PDF Association’s accessible PDF guidance offers useful context for why the intended logical order matters even when the visual page looks clear.
Scanned PDFs need OCR before CAT processing¶
A scanned PDF stores a page as an image rather than as selectable, searchable text. A CAT tool that needs text can’t translate the words directly from that image. OCR, or optical character recognition, first identifies characters in the image and turns them into text that software can process.
OCR doesn’t restore the original document’s full meaning or structure by itself. Recognition can misread characters, and image-based text doesn’t automatically become a properly structured table. Adobe’s PDF accessibility overview explains that OCR converts image characters into text but that the resulting text should be checked for suspect recognition.
A reviewer should compare OCR output with the image, not only scan for obvious spelling mistakes. A misread number, symbol, or label in a table can change the meaning of a row. A missing line can leave a value without its heading. Column order can still be wrong even when every character has been recognized correctly.
The source quality affects how useful that comparison is. A clear printed page provides stronger evidence than a degraded or uneven scan. Handwritten notes, obscured text, and unclear marks can require human review against the image; treating uncertain OCR as confirmed source text creates avoidable risk.
For a scanned table, check the recognized text and the structure as separate items. First confirm that words and values were read correctly. Then confirm that rows, columns, headers, and notes still relate to each other. OCR can produce readable text without recovering those relationships.
Typical example: a scanned report contains a table with two columns of measurements and a footnote below it. OCR may identify most of the words and numbers but place the footnote between rows, or separate a column heading from the values beneath it. The translator needs the original image and a repaired, checked text version to know which content belongs together.
OCR also changes what the CAT tool receives. The tool processes the recognized text, not a perfect interpretation of the scanned page. That distinction is why checking the extraction before translation is more useful than waiting until the final layout review to find a misread label.
For more on how translators check work across stages, see translation QA and the TEP model. When the source itself comes from a scan, source verification is part of that quality process, not a cosmetic extra.
A safer workflow from source file to translated deliverable¶
The most reliable starting point is the original editable file used to create the PDF. A DOCX or another source document usually contains more editable structure than the PDF produced from it. Starting there can avoid some reconstruction problems introduced by converting the PDF back into an editable format.
When the original isn’t available, treat conversion as a review stage rather than a one-click formality. The translator should see the file that will enter the CAT tool, and the team should compare that file against the source PDF before translation begins.
A workflow that makes the handoffs clear looks like this:
- Confirm the source. Ask whether the client can provide the original editable file. If so, check that it matches the supplied PDF and use it as the translation source where appropriate.
- Agree on the deliverable. Specify whether the result should be translated text, an editable document, or a PDF with reconstructed layout. Agree on whether visual reconstruction is part of the work.
- Identify image-only pages. Check whether text can be selected and searched. If the page is image-based, run OCR before text-based CAT processing.
- Convert the file. Choose an intermediate format that suits the task. Don’t assume that a visually faithful conversion also recovers logical reading order.
- Inspect the extraction. Check columns, table relationships, headings, footnotes, and recurring page elements against the PDF. Repair problems in the intermediate file before translation.
- Translate the checked content. Work from segments whose context and sequence make sense. Flag any source content that remains uncertain rather than silently guessing.
- Review the translated layout. Compare the deliverable with the agreed output scope. Check table cell associations and multi-column reading order again because translated text may take up a different amount of space.
A ProZ forum discussion about PDF alignment lists practical conversion cleanup such as removing headers and footers, fixing hard line breaks and hyphenation, repairing columns and tables, and checking spaces around ligatures. The post is an older, anecdotal account rather than a current software specification, but the types of cleanup it describes remain useful preflight checks.
A common failure point: a team checks whether the converted file opens in the CAT tool, but not whether it reads correctly. Opening successfully proves that the file can be processed; it doesn’t prove that the segment sequence or table relationships are right. A brief comparison with the PDF before translation can catch errors that would otherwise spread through the translated file.
Post-translation review needs to match the promised deliverable. If the task is translation of text only, a full page reconstruction may fall outside the agreed scope. If the client needs a formatted file, someone needs to inspect the actual output, not just the CAT segments. A translated table can have accurate wording and still place a value under the wrong header.
For more on the CAT environment around these decisions, compare CAT tools for translators. Tool choice can affect available processing options, but no CAT tool can make an unstructured page’s visual appearance equivalent to reliable semantic structure without checking the result.
Common mistakes that make PDF table problems worse¶
The fastest way to lose time is to assume that a clean-looking PDF will convert cleanly. Visual polish doesn’t tell you whether the file contains tags, whether a table has semantic cell relationships, or whether text is selectable. Check the material the CAT tool will actually receive.
Another common mistake is leaving layout decisions until after translation. A malformed table may be harder to interpret once the segments have been separated and translated. Repairing the source structure first gives the translator clearer context and gives reviewers a more dependable basis for comparison.
Avoid these shortcuts:
- Importing plain text for a layout-sensitive job. Plain text drops formatting and can remove cues that help identify columns, cells, headings, and footnotes.
- Assuming OCR output is accurate because it looks readable. Check every uncertain word, symbol, and value against the scan, then check the relationships between table elements.
- Reviewing a table only for spelling. A terminology check won’t reveal that a value has moved beneath the wrong header.
- Treating a page break as the end of a table. A continuation may repeat its header or carry on an earlier set of rows.
- Confusing appearance with reading order. A page can look normal while its extracted text moves across columns in the wrong sequence.
- Promising a PDF without agreeing on reconstruction. The CAT tool may return an editable intermediate file or plain text, so the expected output needs to be explicit.
When a conversion does fail, don’t patch individual segments until the sequence and structure are understood. Recheck the original page, correct the intermediate document, and then confirm that the corrected version presents the content in the right order. Repairing a few visible mistakes without checking the rest can leave less obvious errors untouched.
Uncertain source content calls for a clear decision. If OCR can’t distinguish a character, or if a table’s layout doesn’t show which label belongs to which value, record the uncertainty and ask for clarification where possible. Guessing may create a translation that reads smoothly but conveys the wrong information.
A sensible handoff records the known limits: which pages were scanned, which tables needed repair, whether footnotes or columns required interpretation, and what format the final review covers. That gives the next person concrete questions to check instead of leaving them to infer what happened during conversion.
If your workflow also includes machine translation, integrating machine translation into a CAT pipeline can help clarify where human checks belong. Regardless of the translation engine, extraction order and source structure remain separate things to verify.
FAQ¶
Why do CAT tools scramble tables when translating a PDF?¶
PDFs can store text by visual position rather than as a semantic table with defined rows, columns, and cell relationships. Conversion may extract the cells in an unexpected order or lose some content, so compare the editable intermediate file with the PDF before translation.
How do I translate a two-column PDF without mixing up the reading order?¶
Check the extracted text against the page and confirm that each column reads from top to bottom in the intended sequence. Repair an interleaved intermediate document before importing it into the CAT tool, and check sidebars and footnotes separately.
Should I convert a PDF to Word before importing it into a CAT tool?¶
Use the original editable source file when it’s available and matches the PDF. If the PDF is the only source, convert it to an editable format, inspect the recovered text and layout, and agree on whether the final deliverable is an editable document or a reconstructed PDF.
Can a CAT tool translate a scanned PDF with tables?¶
A scanned PDF needs OCR first because its page content is stored as images rather than selectable text. Check recognized characters against the scan, then verify that each table value still belongs to the right row and header before translation.
How can I check whether PDF text extraction is in the right order before translation?¶
Compare the extracted text with the page. Check column sequence, table rows and headers, headings, footnotes, and image-only pages; for a tagged PDF, check whether its structure reflects the intended reading order.