Trados Scanned PDF Import: Why It Fails and What Agencies Do

Learn why Trados and memoQ struggle with image-only PDFs, how to diagnose missing text, and how agencies prepare OCR output for CAT tools.

Also in: EN UK RU
Trados Scanned PDF Import: Why It Fails and What Agencies Do

A project arrives as a PDF, but the CAT tool imports a blank document, a few stray symbols, or nothing at all. The file opens normally in a PDF reader, so the failure looks mysterious. Often, the answer is simple: the pages are pictures, not text.

That distinction changes the job. Trados Studio and memoQ can process text in a PDF or a converted file, but an image-only scan first needs optical character recognition, or OCR. OCR software detects characters in page images and turns them into machine-readable text. The result still needs review: recognition can introduce wrong characters, broken reading order, or formatting that doesn’t match the original.

For agencies, the practical question isn’t just “Can the CAT tool open this PDF?” It’s “What source is actually inside the file, how reliable is the extracted text, and what needs checking before translation?” This guide explains how to answer those questions and how to build a workflow that catches conversion problems before they reach a translator or client.

Why a scanned PDF isn’t ordinary text

A PDF is a container for page content. Some PDFs contain selectable text. Others contain only images of pages. The second kind can look just like a digital document on screen, but the visible letters are pixels rather than text characters.

That difference matters because a CAT tool needs machine-readable content to create segments (the individual text units a translator works on). When a PDF has no text layer, the CAT tool has no words to extract. Importing the file again, changing a project setting, or trying another segmentation rule won’t turn page pixels into text. OCR has to happen first.

RWS describes this limitation in its product note about the PDF conversion technology introduced on 13 March 2023 and used from Trados Studio 2022 CU6 onward: the new conversion process won’t convert a merely scanned picture, though it can attempt conversion when scanned text is selectable. RWS’s product note explains why a file can open in a PDF reader yet fail during CAT import.

“PDF should always be seen as a workaround to the situation where the original file cannot be made available.”

— Trados Product Management (RWS)

The point isn’t that every PDF is unusable. A PDF with a good text layer can be a workable input. The warning is about treating a fixed-layout delivery format as if it were a clean authoring file. PDF often captures appearance rather than the document structure that localization tools need.

How to tell whether a PDF has a text layer

Run a quick test before opening a project:

  1. Open the PDF in a reader.
  2. Try to select a word with the cursor.
  3. Copy the word and paste it into a plain text field.
  4. Check whether the pasted text is correct and appears in the expected order.

RWS identifies selectable text as a sign that its newer conversion technology can attempt conversion. If you can’t select text, the PDF is likely image-only and needs OCR before CAT processing. A selection test is a useful triage step, not a quality guarantee: selectable text can still be garbled, incomplete, or out of reading order.

A scan can also contain both kinds of content. For example, one page may have a text layer while another is an image. A PDF may also have OCR text that was added earlier but is inaccurate. Check a sample from each page type rather than assuming the whole file behaves consistently.

This is especially helpful when an import appears to work. A CAT tool can produce a document even if the extracted sequence is wrong. A two-column page might become alternating fragments; a table may be read across rows in the wrong order; a footnote may appear in the middle of a paragraph. The project exists, but the source isn’t ready for translation.

Agencies handling a mix of scans and digital files can use a simple intake distinction: “selectable and coherent,” “selectable but disordered,” or “not selectable.” The labels guide the next step without implying that a PDF is good just because it contains selectable characters.

What the CAT tool sees

A translator sees the page as a layout. A CAT tool sees extracted text and file structure. When a tool converts PDF content to DOCX, it has to reconstruct a document from positioned text, lines, images, and page elements. That reconstruction can be imperfect even when the original text is selectable.

memoQ makes the distinction explicit in its current PDF documentation: it can import PDF content as DOCX-converted content or plain text, but it doesn’t extract text from scanned PDFs whose pages are saved as images. The documentation says to run OCR in an external page-reader program first.

The same guide warns that PDF isn’t a text format and that text flow may be missing or out of order after import. “Text flow” means the sequence in which the content reads, such as heading, paragraph, list, then caption. A page can look correct while its extracted flow is wrong.

For an agency, those details affect more than appearance. Missing text can become an omission. Incorrect reading order can change meaning. A table converted into disconnected text may hide which value belongs to which label. The conversion stage is therefore part of source preparation, not a button the project manager can click and forget.

Teams planning the wider tool choice can compare Trados Studio and memoQ for agency work, but the core issue here comes before the CAT-tool comparison: image pixels must become text before either tool can translate them.

How agencies diagnose import failures

A good diagnosis separates file problems from conversion problems. If the team knows whether the PDF contains text, whether the file is protected, and where the conversion breaks down, it can choose a suitable route instead of repeating the same import attempt.

Start with the original file

Ask the client for the native source file first. That might be a DOCX, a presentation file, or another editable document used to create the PDF. RWS recommends using the original source whenever possible because PDF is a workaround that adds processing effort and doesn’t readily lend itself to localization.

A native file can preserve paragraph structure, tables, headings, and other elements in a way that a PDF reconstruction may not. It can also avoid OCR entirely if the original contains actual text. When the client has both the source file and a scan, ask which one reflects the approved version. An older editable document may not match the signed or finalized PDF.

Requesting the source file doesn’t mean the agency should ignore the PDF. The PDF can remain the visual reference for checking page breaks, labels, stamps, signatures, and other elements that the source file may not show in the same way. The source file is usually a better translation input; the PDF can help confirm what the final document is meant to look like.

Check for access and basic file issues

memoQ’s documentation says password-protected PDFs can’t be imported. If the file requires a password, ask the client for an unlocked copy or for an editable source file rather than treating the access barrier as an OCR problem.

A failed import can also involve a file that is damaged or behaves unexpectedly in a reader. Open the PDF and inspect a few pages before starting conversion. Confirm that the pages render, that the right document was supplied, and that the pages appear to be in the intended order. Those checks won’t fix OCR, but they can prevent the team from diagnosing the wrong issue.

Inspect text quality, not just import status

If a word can be selected, copy it into a plain text field and read it. Look for missing accents, substituted characters, repeated words, or lines that jump between columns. Compare a few extracted passages with the visual page.

Trados guidance says OCR output should be checked and corrected for source errors such as confused characters or names. The same principle applies before importing into memoQ: poor recognition shouldn’t be passed downstream as if it were clean source text. A spelling error in a personal name or a code may be much more consequential than a cosmetic line break.

For an intake team, a quick test page can be more useful than a broad judgment about the file. Pick a page with the densest content, a table, or a section where the scan looks faint. If OCR handles that page poorly, the team has evidence to discuss the next step with the client before estimating the translation work.

Look for a prompt when conversion appears stuck

A stalled conversion isn’t always an OCR failure. In a community discussion about Trados 2024, an RWS staff response described a pause related to Microsoft Word opening and displaying a confirmation prompt. The user reported waiting several minutes before noticing the prompt. This is an anecdotal lead, not a universal fix, but checking for a hidden Word dialog can help when a conversion process appears to stop without an error.

The Trados Community discussion also notes that the newer conversion route differs from the older one. A delay may be a separate issue from the missing text layer, so identify what the software is doing before changing the workflow.

The agency workflow: OCR first, CAT second

The safest general workflow separates recognition, correction, translation, and final layout. Each stage has a different goal. OCR turns the image into text; review checks that text; the CAT tool supports translation; layout work rebuilds the deliverable where needed.

Step 1: choose the best source

Ask for an editable original before processing the scan. If the client has no native file, confirm whether the PDF contains selectable text. If it doesn’t, tell the client that OCR is required and that conversion quality will affect the amount of review and layout work.

The client should also know what the agency can and can’t infer from a scan. A blurry stamp, handwritten note, or faint annotation may be visible but not reliably machine-readable. Don’t silently guess unclear content. Mark uncertain text for human review or ask the client to confirm it.

Step 2: run OCR or convert to DOCX

For an image-only PDF, use a page-reader or OCR program to recognize the text and save a DOCX. memoQ’s current guidance names Nuance OmniPage and ABBYY FineReader as examples of external programs for this task. RWS’s March 2023 product announcement also lists opening PDFs, including OCRed PDFs, in Microsoft Word and saving as DOCX; using Adobe’s PDF-to-Word option; or using a third-party PDF or OCR application such as ABBYY FineReader or Readiris.

Those are options described in the cited vendor guidance, not a guarantee that any specific program will produce a clean file for every scan. Product behavior and availability can change. The agency should test the actual document and tool combination instead of assuming one route always wins.

RWS described PDF Assistant for Trados Studio as another option for improved scanned-PDF conversion. In its first release, the add-in used Microsoft Word behind the scenes to convert PDFs to DOCX and relied on Word for scanned documents, bidirectional scripts, and Asian languages. That description refers to the first release, so check current compatibility and behavior before making it part of a production workflow. RWS’s PDF Assistant article also cautions that conversion isn’t guaranteed to be fully accurate and that specialist PDF editing software may be needed when an add-in can’t process a file correctly.

The older IRIS route is another reason to check version-specific advice before following a tutorial. RWS says it discontinued the older PDF file type that worked with IRIS starting in March 2023, and its IRIS AppStore page states that Studio 2022 version 4.0.0.0 onward is incompatible with IRIS. Advice written for an older setup may not describe a usable current workflow.

Step 3: compare the DOCX with the PDF

Don’t import the converted file until someone has inspected it. Compare the DOCX with the PDF page by page, focusing first on the places where automated conversion commonly causes visible differences:

  • headings and paragraphs that should be separate;
  • columns, tables, and labels;
  • footnotes and captions;
  • repeated headers or footers;
  • symbols, names, and special characters;
  • text that appears faint, skewed, or crowded on the scan.

The goal isn’t to reproduce every page detail before translation. The goal is to confirm that the text is complete and in a sensible reading order. Correcting a broken paragraph or a table label now is easier than discovering the problem in a translated file.

Step 4: select the CAT import mode

memoQ’s PDF conversion offers a trade-off. A setting that preserves text flow can prioritize a readable sequence while losing some formatting. A setting that tries to retain formatting can risk losing some text. The right choice depends on whether the project needs reliable text order, closer visual resemblance, or both through separate review.

Plain-text import discards formatting. memoQ’s documentation doesn’t recommend that route for documents being translated, though it can serve LiveDocs corpus import or alignment. For a standard translation job, an unformatted text dump may make it harder to distinguish headings, columns, labels, and table relationships.

Trados Studio 2024 SR1 documentation describes three PDF layout recovery modes: Flowing, Continuous, and Exact. Exact uses text boxes in Word to recover the page presentation; Flowing preserves text flow and page elements. Those settings are version-specific and should be checked against the installed Studio version in the Trados PDF file type documentation.

A practical choice is to decide what the translator needs to work safely. If the source is a form where the association between a prompt and its answer matters, a readable structure can be more useful than visual similarity. If the translator needs page-level context, preserving the layout may help, but the team still needs to verify that no text disappeared.

Step 5: preview the conversion before committing

Trados guidance describes previewing the expected conversion and saving a DOCX for inspection before starting translation. That gives the agency a chance to catch errors in the source preparation stage instead of finding them after segments have been translated.

RWS also warns that the new PDF processing mechanism can differ from the older file type in analysis statistics, translation-unit lookup, image recovery, special-character handling, and PerfectMatch transfer. For Asian or other non-Latin source languages, RWS suggests enabling alternative processing. The note is tied to the processing changes introduced from March 2023 and Studio CU6 onward, so treat it as release-specific guidance rather than a promise about every installation.

Step 6: translate and prepare the final deliverable

Once the DOCX has passed source checks, create the CAT project, translate, and review the target against both the converted source and the original PDF. The translated DOCX may not reproduce the final page design. If the client expects a publication-ready PDF, agree who will rebuild or adjust the layout and which file will be the visual reference.

memoQ’s documentation says it doesn’t export a translated PDF from a PDF source. Depending on the import method, memoQ exports plain text or DOCX. The agency therefore needs a separate step to create the final PDF when the requested deliverable requires one. That step may involve a layout specialist or another editing tool, depending on the document.

Agencies new to this kind of work can also review how to quote a scanned document when word count finds nothing. A scan with no extractable text changes the estimating process: the team must account for recognition and checking, not just translation.

What goes wrong after OCR

OCR output is a draft of the source text, not a verified transcript. A successful conversion means that the software produced a file. It doesn’t prove that every word, number, symbol, or relationship between page elements survived.

“Can’t import scanned PDF files: memoQ doesn’t extract text from scanned PDF files, where the pages are saved as images and not as text.”

— memoQ Help documentation

The distinction is useful when explaining a failed import to a client. The issue isn’t that the CAT tool has ignored readable content. The page image needs OCR before the software can process the words.

Recognition errors

OCR can confuse characters that look similar, especially when the source is faint or crowded. Names and unusual terms deserve particular attention because a spellchecker may not flag a wrong character as an error. The Trados PDF guidance identifies scan quality and language as factors in OCR accuracy and describes skewed, blurry, faint, smudged, and handwritten text as poor candidates for the workflow discussed there.

Don’t treat that list as a universal test for every OCR product. The cited Trados article describes an older Solid Documents OCR workflow. Its practical warning still supports a sensible intake rule: inspect difficult pages closely, and don’t assume that a machine-readable result is an accurate one.

Broken reading order

A PDF can place words at precise positions on a page without storing them in the order a person reads them. A conversion process may then capture fragments from separate columns or boxes in the wrong sequence. The output might contain all the words and still be hard to interpret.

memoQ’s warning that text flow may be missing or out of order applies directly to this problem. When an import produces an odd sequence, compare the segment order with the visual page. Correcting the source DOCX may be more reliable than asking the translator to infer the intended order from disconnected fragments.

Lost or reconstructed formatting

A table can become plain lines; text boxes can split a sentence; page headers can enter the main text. Layout recovery modes offer different compromises rather than a guaranteed copy of the original. An agency should tell the client whether the translation stage will preserve content structure, visual design, or both through separate work.

Trados Studio 2024 SR1’s Flowing, Continuous, and Exact modes illustrate that trade-off. Exact uses Word text boxes to recover page presentation, while Flowing prioritizes text flow and page elements. Because these descriptions are specific to that version, verify the settings in the installed release before building a procedure around them.

Characters and scripts that need extra care

RWS says the new PDF file type can differ from the older one in special-character handling and recommends alternative processing for Asian or other non-Latin source languages. That doesn’t mean every non-Latin PDF will fail. It means the agency should test the relevant language and inspect the converted result rather than assume that Latin-script behavior carries over.

A Trados Community reply about the newer processing path said:

“The new one isn’t as functional, hence the reason why we built the plugin to try and provide better support.”

— Paul Filkin, RWS, in the Trados Community

The comment refers to the discussion’s specific conversion context, not every current Trados workflow. For an agency, the useful takeaway is to validate its actual setup with representative files and avoid assuming that behavior from an older PDF filter will carry over unchanged.

Old advice and unavailable routes

Search results and internal SOPs can keep pointing teams toward older PDF workflows. The IRIS notice says the legacy integration became incompatible with Studio 2022 version 4.0.0.0 onward. memoQ’s current documentation also says its TransPDF service isn’t available in memoQ 10.4 and later. A tutorial that relies on either route may describe a historical process, not a current option.

Historical information can still explain why a team remembers a different workflow. A memoQ blog post about PDFs in memoQ 8.1 describes the former TransPDF integration and notes that OCR could introduce character and formatting problems. The post isn’t a current workflow recommendation; memoQ’s current documentation says TransPDF is unavailable from memoQ 10.4 onward.

Preventable mistakes in an agency workflow

Scanned-PDF issues often become expensive because the team discovers them after work has started. A consistent intake and quality check can catch the main risks early.

Treating “opened” as “ready”

A PDF opening in a reader only proves that the reader can display it. It doesn’t prove that the text is selectable, complete, or in reading order. Run the selection test, copy a sample, and inspect a difficult page before creating the project.

Skipping source-file requests

Agencies sometimes accept the PDF without asking whether an editable source exists. RWS recommends using the original file where possible because PDF adds conversion effort and isn’t designed for localization. A short request at intake can avoid unnecessary OCR and repair.

If the client has no source file, set expectations before quoting or scheduling. Explain that OCR and source verification are separate from translation, and that poor scans may need more manual checking. Don’t promise exact formatting based only on the fact that the PDF looks clear on screen.

Assuming OCR fixed every page

Mixed files may contain both selectable text and image-only pages. A conversion can handle one page well and fail on another. Inspect representative pages from the beginning, middle, and end, plus any page with tables, columns, forms, or unusual scripts.

Choosing a layout setting without considering purpose

A visually close conversion isn’t automatically the best translation source. If text flow matters more than some formatting, memoQ’s documentation says to favor flow preservation. If formatting matters more, its conversion can attempt to keep it while accepting that some text may be lost. In either case, review the result against the source.

Importing plain text for a formatted translation

Plain-text import loses formatting and memoQ doesn’t recommend it for documents being translated. It may suit corpus import or alignment, but it can remove the context that makes labels, tables, and page elements understandable to a translator. Pick it only when the project needs plain text rather than a structured translation file.

Forgetting the output format

A client who supplies a PDF may expect a translated PDF back. memoQ doesn’t export a translated PDF from a PDF source; it exports plain text or DOCX, depending on the import method. Confirm the target format during intake and plan a separate layout step when the final deliverable must be PDF.

Teams that need to weigh CAT-tool behavior against wider project needs can read a freelancer-focused Trados and memoQ comparison. For scanned files, though, neither tool removes the need to verify OCR output before translation.

FAQ

Why does Trados Studio fail to import a scanned PDF?

An image-only PDF contains page pictures rather than machine-readable text. RWS says its newer PDF conversion technology won’t convert a merely scanned picture, so OCR must recognize the text and create a usable text layer or editable file first.

Can memoQ import a scanned PDF directly?

No. memoQ’s current documentation says memoQ doesn’t extract text from scanned PDFs whose pages are saved as images. Run OCR in an external page-reader program, save an editable file such as DOCX, and import the reviewed result.

How do agencies convert scanned PDFs for Trados or memoQ?

Agencies run OCR or PDF-to-DOCX conversion, inspect the text and layout against the original, correct recognition errors, and then import the checked file into the CAT tool. RWS lists Microsoft Word, Adobe PDF-to-Word, and third-party tools such as ABBYY FineReader and Readiris as options in its March 2023 product announcement.

What should an agency do when OCR scrambles a PDF layout?

Compare the converted file with the original page and correct the reading order, missing text, and important formatting before translation. Choose a conversion approach according to whether text flow or visual layout matters more, and request the original editable file if the PDF remains difficult to process.

Does Trados PDF Assistant OCR scanned documents?

RWS described the first release of PDF Assistant as using Microsoft Word behind the scenes, including Word for scanned documents. The description applies to that first release, so check current compatibility and behavior before relying on it for a production job.

Can memoQ export a translated PDF from a PDF source?

No. memoQ’s documentation says it exports plain text or DOCX, depending on how the PDF was imported. If the client needs a final PDF, plan a separate layout or publishing step after translation.

Try ChatsControl

AI platform for professional translators

Try for free →