PdfParse Docs

Zapier: send PDFs in and extracted rows out

Build a two-Zap invoice workflow: upload PDFs to PdfParse, then send saved invoice fields to Google Sheets without writing code.

Connect a PdfParse project to Zapier to move documents into an extraction table and use its saved rows in another app. This guide builds one concrete workflow: invoices arriving in a Google Drive folder become rows in a Google Sheet.

Availability: the native PdfParse app is being prepared for release. These instructions describe the integration in testing; they become usable in the Zap editor once you have access to its private invite or public listing. If PdfParse is absent from the app picker, do not choose another app with a similar name. Check the integration page for availability.

What you will build

Use two Zaps because document extraction runs in the background:

  1. Document intake: Google Drive → PdfParse Upload PDF and Start Extraction.
  2. Saved data: PdfParse New Extracted Row → Google Sheets Create Spreadsheet Row.

The first Zap returns a document ID and job ID as soon as extraction is queued. The second reads rows after they have been saved. You do not need to guess how long OCR takes or add a fixed delay.

For the example, use the synthetic invoice PDF, which contains invoice HOS-2026-1042 from Harbor Office Supply, with currency USD and total 70.20. The amount is an example, not an extraction accuracy guarantee.

Prepare the project, table, and spreadsheet

You need a PdfParse account with an active subscription and enough extraction allowance, a Google Drive folder, a Google Sheet, and a Zapier account that supports the triggers and actions you select. Plan requirements and polling schedules depend on Zapier; check them before turning on your Zaps.

  1. Open your PdfParse project, or follow Getting started to create one. Use a clear project name such as Accounts Payable.
  2. Create an Invoices table with the fields below. If you already use an invoice table, reuse it and adapt the mapping to its field names.
  3. Extract the sample invoice in PdfParse first. Inspect the saved record alongside the source PDF before automating downstream work.
  4. Create a Google Sheet with these headers in the first row: Invoice number, Supplier, Currency, Total, and Source document ID.
PdfParse fieldTypeExtraction instruction
invoice_numberTextExtract the invoice identifier printed on the document.
supplierTextExtract the supplier's business name.
currencyTextExtract the three-letter currency code.
totalNumberExtract the final invoice total including tax.

Keep the table's automatic integer id column. New Extracted Row uses it to identify and order records. Do not replace it with an invoice number or an extracted business identifier.

Check the result: the sample record should contain HOS-2026-1042, Harbor Office Supply, USD, and numeric 70.2. A spreadsheet can display the amount as 70.20. If a field is missing or incorrect, adjust the extraction instructions and review the PDF before proceeding.

Connect the right PdfParse project

Each Zapier connection uses a project API key. A connection can access the project that owns the key; typing a different connection name does not change its access.

  1. In the project's navigation, open API Keys and select Create API Key.
  2. Label the key Zapier — Accounts Payable. Choose an expiration appropriate for the workflow and save the key securely when it is shown.
  3. In the Zap editor, select PdfParse for the relevant step, then connect an account.
  4. Enter the key in Project API Key and enter Accounts Payable in Connection Name.
  5. Test the connection. It checks that the key can list the project's tables.

The test can succeed when the project has no tables. If Invoices does not appear in a step's Table list, create it in that project and refresh the fields. To connect another project, create a separate connection with that project's key.

Never place the key in a spreadsheet, file URL, or ordinary text field in the Zap. Reconnect when the key expires, is revoked, or is replaced. See Authentication for the underlying access rules.

Build the PDF intake Zap

This Zap sends a downloadable PDF to the existing Invoices table.

  1. Create a Zap with Google Drive as the trigger app. Choose the trigger for a new file in a folder and select your invoice intake folder.
  2. Test the Drive trigger using the sample PDF. If the folder also receives non-PDF files, add a filter that allows only PDF filenames or PDF MIME types.
  3. Add a PdfParse action and choose Upload PDF and Start Extraction.
  4. Select the Accounts Payable connection and the Invoices table.
  5. For PDF File, map the Drive trigger's downloadable file. Use the file output rather than a sharing link or browser preview URL.
  6. For PDF Filename, map the source filename, including its .pdf extension.
  7. Test the action once.

Expected result: the action returns job_id, document_id, status, table_id, and filename. The initial status is normally queued. The action deliberately returns identifiers rather than waiting for the invoice fields.

Verify it: find the uploaded PDF in the PdfParse project and inspect its processing state. If needed, use Find Extraction Job with the returned job_id to inspect the current job. Wait until a saved invoice record is visible before testing the next Zap.

Recover: a PDF must be downloadable over HTTPS and at most 8 MiB for this action. A sharing page, login page, image, or renamed non-PDF file is rejected. If the action fails after uploading, check whether the document or job already exists before replaying the entire action. A full replay can create another upload and another record; the job request is protected against duplicate requests for the same upload, but the whole Zap action is not guaranteed to run only once.

Build the saved-row Zap

This Zap reads the extracted fields and writes one spreadsheet row per saved record.

  1. Create another Zap with PdfParse as the trigger app.
  2. Choose New Extracted Row, the Accounts Payable connection, and the Invoices table.
  3. Test the trigger. Select the invoice record you reviewed earlier.
  4. Add Google Sheets and choose Create Spreadsheet Row. Connect your Google account and select the spreadsheet and worksheet prepared above.
  5. Map the fields using this table. The invoice fields appear under Extracted in the PdfParse output.
  6. Test the Sheets action and compare the new spreadsheet row with the source invoice.
Google Sheet headerPdfParse outputExpected sample value
Invoice numberExtracted → invoice_numberHOS-2026-1042
SupplierExtracted → supplierHarbor Office Supply
CurrencyExtracted → currencyUSD
TotalExtracted → total70.2
Source document IDSource Document IDThe uploaded document's ID

Verify it: the spreadsheet row should agree with the reviewed PdfParse record. Preserve Source Document ID so you can trace a spreadsheet entry back to its input. Zapier Row ID identifies the table row and can be used in your own duplicate-prevention logic.

If you add or rename extraction fields later, refresh the trigger's fields and review the spreadsheet mapping before restarting the workflow. For repeating invoice items, use the child table as a separate row trigger; its rows are separate records and may require a different worksheet. Find Extracted Row returns the first matching record, not every line item belonging to a document.

Turn on the workflow and verify a fresh document

  1. Turn on the saved-row Zap first, then the intake Zap. Testing a Zap can write real rows into the chosen worksheet, so identify and remove test entries as appropriate.
  2. Add a new PDF with a different invoice number to the watched Drive folder.
  3. Check the intake Zap's history for its job and document IDs.
  4. Inspect the extracted record in PdfParse. After the next scheduled poll, confirm that the saved-row Zap creates the corresponding spreadsheet row.
  5. Compare the invoice number, supplier, currency, total, and source ID across the document, PdfParse, and spreadsheet.

The automation now has a verifiable path from PDF to saved record to spreadsheet. The saved-row trigger starts from new rows; it does not backfill the entire table or run again when you edit an existing record.

Understand polling and request limits

New Extracted Row reads the newest 200 rows ordered by automatic integer id on each poll. If more than 200 rows arrive between polls, older unseen rows can fall outside that window. It is intended for moderate-volume tables, not a bulk synchronization pipeline.

Polling frequency depends on your Zapier plan. PdfParse project API keys currently allow 1,000 requests per UTC day, shared by all uses of that key. A five-minute polling schedule makes about 288 requests per day per enabled row Zap; setup, uploads, other Zaps, and lookups add requests. One-minute polling alone would exceed the daily allowance. Use a schedule and number of enabled Zaps that fit the allowance. Do not use multiple keys simply to bypass it.

For high-volume or auditable bulk delivery, use the REST API with pagination and a durable consumer. Do not depend on this polling trigger for a guaranteed full history.

Add completion or failure notifications

The integration also has two instant job triggers: Extraction Completed and Extraction Failed. These use signed project webhooks and create or remove their subscriptions when the Zap is turned on or off.

To monitor failures, create a Zap with Extraction Failed, connect the project, and send the resulting job ID and status to your chosen notification app. Test it using a genuine failed job in that project. If no failed jobs exist yet, the test list can be empty; the sample schema is not proof of a live failure.

Extraction Completed reports that a job has finished. A multi-document job can complete with some documents unsuccessful. Neither job trigger provides invoice fields or approves the data for payment. Use New Extracted Row for saved-data delivery, and review documents and amounts before any financial action.

Troubleshoot the first run

SymptomWhat to check
PdfParse is missing from the app pickerAccess to the private integration invite or the public listing. The app must be registered and released before native steps are usable.
Connection failsThe project key is valid, unexpired, unrevoked, and can read tables.
Wrong tables appearThe connection uses a key from a different project. The connection name is only a label.
File action rejects the inputMap the downloadable PDF, not the preview page; check the .pdf filename and 8 MiB limit.
Upload succeeds but extraction failsCheck the active subscription, extraction allowance, table selection, and the document's processing details.
No trigger test recordsSave at least one record in the selected table, then retest.
Some spreadsheet rows are absentCheck Zap history, polling schedule, the 200-row window, and the API key's daily request limit.
Duplicate spreadsheet rowsCheck test runs, manual replays, and repeat uploads. Add downstream lookup logic using the row ID where needed.
Job trigger stops after reconnectingTurn the Zap off and on again to refresh the subscription under the intended project connection.

For a programmatic workflow with explicit retries and pagination, follow Run an extraction job. For this workflow, the next step is to run one fresh invoice through both Zaps and verify the resulting spreadsheet row.