Extract Text and Tables from PDF Files
Extract clean text and structured table data from any PDF. Perfect for data pipelines, search indexing, content analysis, and feeding AI/LLM models.
# JSON with tables — /v2/pdf/extract-text (recommended)
curl -X POST \
https://api.convertfilefast.com/v2/pdf/extract-text \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@report.pdf" \
-F "pages=1-5" \
-F "extract_tables=true"
# Plain text file — /v2/convert/pdf-to-txt
curl -X POST \
https://api.convertfilefast.com/v2/convert/pdf-to-txt \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@report.pdf" \
--output extracted.txtimport requests
# Option 1: JSON with tables (extract-text) — rich structured output
response = requests.post(
"https://api.convertfilefast.com/v2/pdf/extract-text",
headers={"Authorization": "Bearer YOUR_API_KEY"},
files={"file": open("report.pdf", "rb")},
data={"pages": "1-5", "extract_tables": "true"}
)
data = response.json()
for page in data["pages"]:
print(f"Page {page['page_number']}: {page['text'][:100]}...")
# Option 2: Plain text file (pdf-to-txt) — simple .txt output
response = requests.post(
"https://api.convertfilefast.com/v2/convert/pdf-to-txt",
headers={"Authorization": "Bearer YOUR_API_KEY"},
files={"file": open("report.pdf", "rb")}
)
with open("extracted.txt", "wb") as f:
f.write(response.content)const formData = new FormData();
formData.append("file", fs.createReadStream("report.pdf"));
formData.append("pages", "1-5");
formData.append("extract_tables", "true");
const response = await fetch(
"https://api.convertfilefast.com/v2/pdf/extract-text",
{
method: "POST",
headers: { "Authorization": "Bearer YOUR_API_KEY" },
body: formData,
}
);
const data = await response.json();
console.log(`Extracted ${data.total_pages} pages`);// n8n HTTP Request Node
{
"method": "POST",
"url": "https://api.convertfilefast.com/v2/pdf/extract-text",
"contentType": "multipart-form-data",
"bodyParameters": {
"parameters": [
{ "name": "file", "parameterType": "formBinaryData", "inputDataFieldName": "data" },
{ "name": "extract_tables", "value": "true" }
]
},
"options": { "response": { "responseFormat": "json" } }
}Advantages
Why use our API?
Complete and reliable solution for integration in any tech stack.
Text Extraction
Extract clean, structured text from any PDF with proper paragraph and sentence detection.
Table Detection
Automatically detects and extracts tables as structured data arrays for processing.
Page Selection
Extract from specific pages (e.g., "1,3,5-7") or all pages in one request.
Metadata Access
Get PDF metadata (title, author, creation date) alongside the extracted text content.
AI/LLM Ready
Perfect for feeding extracted content to ChatGPT, Claude, or custom AI model pipelines.
Pipeline Integration
Easily integrate into ETL pipelines with n8n, Airflow, or custom processing scripts.
Or call this from Claude, Cursor, or any MCP client
Convert File Fast ships a native Model Context Protocol server. Give this conversion to an AI agent as a single tool call — no HTTP client, no binary handling, per-user API keys.
npx -y convertfilefast-mcpconvert_file({ target_format: "txt", source_url: "https://example.com/file.pdf" })Start Extracting Text from PDFs
Get your API key and extract text from PDF documents in seconds. Free plan includes 10 conversions per month.
No credit card. 10 free conversions on Free plan.