How to correctly extract text from a pdf using pdf.js

Tags:

I'm new to ES6 and Promise. I'm trying pdf.js to extract texts from all pages of a pdf file into a string array. And when extraction is done, I want to parse the array somehow. Say pdf file(passed via typedarray correctly) has 4 pages and my code is:

let str = [];
PDFJS.getDocument(typedarray).then(function(pdf) {
  for(let i = 1; i <= pdf.numPages; i++) {
    pdf.getPage(i).then(function(page) {
      page.getTextContent().then(function(textContent) {
        for(let j = 0; j < textContent.items.length; j++) {
          str.push(textContent.items[j].str);
        }
        parse(str);
      });
    });
  }
});

It manages to work, but, of course, the problem is my parse function is called 4 times. I just want to call parse only after all 4-pages-extraction is done.

874

asked Nov 16 '16 15:11

Sangbok Lee

2 Answers

Similar to https://stackoverflow.com/a/40494019/1765767 -- collect page promises using Promise.all and don't forget to chain then's:

function gettext(pdfUrl){
  var pdf = pdfjsLib.getDocument(pdfUrl);
  return pdf.then(function(pdf) { // get all pages text
    var maxPages = pdf.pdfInfo.numPages;
    var countPromises = []; // collecting all page promises
    for (var j = 1; j <= maxPages; j++) {
      var page = pdf.getPage(j);

      var txt = "";
      countPromises.push(page.then(function(page) { // add page promise
        var textContent = page.getTextContent();
        return textContent.then(function(text){ // return content promise
          return text.items.map(function (s) { return s.str; }).join(''); // value page text 
        });
      }));
    }
    // Wait for all pages and join text
    return Promise.all(countPromises).then(function (texts) {
      return texts.join('');
    });
  });
}

// waiting on gettext to finish completion, or error
gettext("https://cdn.mozilla.net/pdfjs/tracemonkey.pdf").then(function (text) {
  alert('parse ' + text);
}, 
function (reason) {
  console.error(reason);
});

<script src="https://npmcdn.com/pdfjs-dist/build/pdf.js"></script>

115

answered Sep 19 '22 13:09

async5

A bit more cleaner version of @async5 and updated according to the latest version of "pdfjs-dist": "^2.0.943"

import PDFJS from "pdfjs-dist";
import PDFJSWorker from "pdfjs-dist/build/pdf.worker.js"; // add this to fit 2.3.0

PDFJS.disableTextLayer = true;
PDFJS.disableWorker = true; // not availaible anymore since 2.3.0 (see imports)

const getPageText = async (pdf: Pdf, pageNo: number) => {
  const page = await pdf.getPage(pageNo);
  const tokenizedText = await page.getTextContent();
  const pageText = tokenizedText.items.map(token => token.str).join("");
  return pageText;
};

/* see example of a PDFSource below */
export const getPDFText = async (source: PDFSource): Promise<string> => {
  Object.assign(window, {pdfjsWorker: PDFJSWorker}); // added to fit 2.3.0
  const pdf: Pdf = await PDFJS.getDocument(source).promise;
  const maxPages = pdf.numPages;
  const pageTextPromises = [];
  for (let pageNo = 1; pageNo <= maxPages; pageNo += 1) {
    pageTextPromises.push(getPageText(pdf, pageNo));
  }
  const pageTexts = await Promise.all(pageTextPromises);
  return pageTexts.join(" ");
};

This is the corresponding typescript declaration file that I have used if anyone needs it.

declare module "pdfjs-dist";

type TokenText = {
  str: string;
};

type PageText = {
  items: TokenText[];
};

type PdfPage = {
  getTextContent: () => Promise<PageText>;
};

type Pdf = {
  numPages: number;
  getPage: (pageNo: number) => Promise<PdfPage>;
};

type PDFSource = Buffer | string;

declare module 'pdfjs-dist/build/pdf.worker.js'; // needed in 2.3.0

Example of how to get a PDFSource from a File with Buffer (from node types) :

file.arrayBuffer().then((ab: ArrayBuffer) => {
  const pdfSource: PDFSource = Buffer.from(ab);
});

answered Sep 19 '22 13:09

Atul

Related questions
                            
                                Custom filter giving "Cannot read property 'slice' of undefined" in AngularJS
                            
                                How do you get the key at specifc index in javascript map object?
                            
                                React, Uncaught RangeError: Maximum call stack size exceeded
                            
                                Safe feature-based way for detecting Google Chrome with Javascript?
                            
                                jQuery focus() sometimes not working in IE8
                            
                                How to do a JavaScript HTML 5 Canvas image "page flip" like you commonly see in Flash?
                            
                                Javascript regex to match a string that ends with some characters but not with a particular combination of those
                            
                                Google maps removing controls
                            
                                Stop Twitter Bootstrap Carousel at the end of it's slides
                            
                                Get parts of URL after domain name //... by splitting URL into array
                            
                                How do I see if two rectangles intersect in JavaScript or pseudocode?
                            
                                Getting a strange error with grunt: Object Gruntfile.js has no method 'flatten'
                            
                                D3.js - why mouseover and mouse out fired for each child elements?
                            
                                Disable horizontal scroll; but allow vertical scroll
                            
                                jQuery change method in vanilla JavaScript?
                            
                                Ionic + Angular POST request return state 404
                            
                                Convert curl GET to javascript fetch
                            
                                jQuery, change-event ONLY when user himself changes an input
                            
                                Http Post in laravel using fetch api giving TokenMismatchException
                            
                                react-router index route always active

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

How to correctly extract text from a pdf using pdf.js

Tags:

javascript

callback

pdf

es6-promise

pdf.js

Sangbok Lee

People also ask

2 Answers

async5

Atul

Recent Activity

Donate For Us