Hybrid (base + metered)
Available
SOC 2 certified
Available
Document Processing
Overview & Positioning
Unstructured is a data preprocessing platform built for teams building retrieval-augmented generation (RAG) pipelines, LLM applications, and other systems that need to turn raw, unstructured documents into structured, machine-readable formats. It's aimed squarely at data engineers and ML engineers working on ingestion pipelines who need to extract, transform, and route content from a mix of file types into vector databases or downstream storage systems.
Pricing
Unstructured offers three tiers: a free "Let's Go" plan (15,000 pages/month), Pay-As-You-Go at $0.03/page beyond that, and a Contact-Sales Business tier for dedicated/VPC deployments. Usage-based overage applies once the free monthly page allowance is exceeded, capping at $3,000/month before scaling back to free.
| Tier | Starting At | Limits / Key Inclusions |
|---|---|---|
| Let's Go (Free) | $0/mo | 15,000 free pages/mo, resets monthly, no card required, all features included |
| Pay-As-You-Go | $0.03/page | After the free 15,000 pages/mo; bill caps at $3,000/mo, then free up to 1M pages/mo |
| Business | Contact Sales | Dedicated instance or VPC, multi-user access, full data isolation, dedicated technical support |
See the official website for complete tier limits, add-ons, and enterprise custom pricing. Visit official pricing →
API Access and Certification
Unstructured exposes API access, so the platform can be wired directly into existing data pipelines instead of depending on a UI-driven workflow. It also holds SOC 2 certification, a point worth noting for organizations with compliance obligations around data handling and security when they're trusting a third party to process potentially sensitive documents.
Core Features
The platform handles more than 50 document and image file types, extending well past plain text or PDF into the range of formats typically found in enterprise document processing. It also ships with 40+ source and destination connectors, so data can move in and out of a wide range of external systems without building custom integrations for each one.
For parsing, Unstructured offers four partitioning strategies — Auto, Fast, High-Res, and VLM (vision-language model) — letting users trade off speed against extraction accuracy based on how complex a document is. Chunking works similarly: there are five strategies to choose from — Character, Title, Page, Similarity, and Contextual — each shaping how extracted content gets segmented before it's used downstream, for instance in embedding generation.
Enrichment features round out the pipeline: metadata extraction, generative OCR, image and table description generation, and named entity recognition (NER). These add structured context on top of raw extracted content, which can pay off in better retrieval quality or more precise downstream filtering.
The platform also generates embeddings across several providers — VoyageAI, Bedrock, Azure OpenAI, IBM, and TogetherAI — so teams can produce vectors within the same pipeline instead of shipping content off to a separate service. That also leaves room to match the embedding provider to whatever infrastructure or model preferences a team already has in place.
Deployment options cover SaaS Cloud-Hosted, Dedicated Instance, In-VPC, and Bare Metal, spanning teams happy with a fully managed cloud setup to those that need processing confined to their own VPC or run on dedicated or bare-metal hardware for tighter control.
Rounding things out, Unstructured includes full ETL orchestration plus data duplicate prevention — so the platform isn't just handling extraction and transformation, it's also covering the pipeline orchestration and deduplication work that teams would otherwise have to build themselves.
Summary
Unstructured pairs a free entry tier with metered overage pricing, API-first access, and SOC 2 certification, backed by a feature set that covers broad file-type support, a large connector library, configurable partitioning and chunking, enrichment options, multi-provider embedding generation, and flexible deployment models. Taken together, it's positioned as infrastructure for teams building document-processing pipelines that feed into LLM or search applications.
Verified Core Features
- 15,000 free pages every month
- Support for 50+ document and image file types
- 40+ source and destination connectors
- Partitioning strategies: Auto, Fast, High-Res, VLM
- Chunking strategies: Character, Title, Page, Similarity, Contextual
- Enrichments: Metadata, Generative OCR, Image/Table description, NER
- Embedding generation across VoyageAI, Bedrock, Azure OpenAI, IBM, TogetherAI
- Deployment options: SaaS Cloud-Hosted, Dedicated Instance, In-VPC, Bare Metal
- Full ETL Orchestration & Data Duplicate Prevention
Strengths & Trade-offs
Strengths
- Free tier available before any purchase commitment.
- Public API for integrating with external systems and pipelines.
- SOC 2 certified, per the vendor's own disclosure.
Tradeoffs
- Hybrid pricing — a flat base plus usage-based charges.