⚡ High-speed Java parser for text extraction, PDF ingestion, StAX streaming XLSX & CSV tables, and FastRegex SIMD whitespace normalization.
FastContentParse extracts text from plain files, Markdown, RTF, PDF documents, CSV sheets, and OpenXML Excel spreadsheets (.xlsx), then normalizes it for embedding and retrieval pipelines. It is designed to work alongside FastContentChunk, FastAIVectorDB, and FastAIRag to accelerate text extraction and Parent-Child context retention.
import fastcontentparse.FastContentParse;
import fastcontentparse.ParsedDocument;
import java.nio.file.Path;
public class Demo {
public static void main(String[] args) throws Exception {
// 1. Initialize Document Parser
FastContentParse parser = new FastContentParse();
// 2. Parse PDF / RTF / Markdown Document
ParsedDocument doc = parser.parseFile(Path.of("docs/sample.pdf"));
// 3. Inspect Extracted Type and Normalized UTF-8 Text
System.out.println("Document Type: " + doc.getType());
System.out.println("Extracted Text Preview:\n" + doc.getText().substring(0, 200) + "...");
}
}- Why FastContentParse?
- Key Features
- Performance Benchmarks
- Architecture Overview
- API Quick Reference
- Technical Demos & Benchmarks
- Installation
- Documentation
- Platform Support
- License
- Related Projects
Most Java content pipelines rely on brittle file readers or heavyweight libraries. FastContentParse is focused on the most common content sources for retrieval workflows: text, Markdown, RTF, and PDF.
It provides:
- Simple file parsing with consistent normalization across all formats.
- Layout-based visual paragraph detection for PDF documents using line Y-coordinate offsets to preserve natural section boundaries.
- PDF extraction via Apache PDFBox without requiring full desktop document frameworks.
- Single-Pass RTF stripper eliminating 4 sequential regex passes.
- Optional native tokenizer integration through the separate
FastContentChunkmodule for SIMD-accelerated chunking.
| Feature | Apache Tika | Standard POI / PDFBox raw | FastContentParse |
|---|---|---|---|
| Paragraph Reconstruction | Basic flat stream | Line-by-line fragmented text | Geometry Y-offset visual clustering |
| RTF Parsing Overhead | Full DOM / RTF EditorKit | Heavy regex passes | Single-pass 0-regex byte stripper |
| Spreadsheet Ingestion | Full DOM memory model | Apache POI massive heap bloat | StAX streaming (ZIP + sharedStrings) |
| Pipeline Integration | Generic metadata objects | Ad-hoc manual text parsing | Direct output for FastContentChunk/RAG |
- 📄 Multi-Format Text Extraction — Extracts clean text from PDF, RTF, Markdown, images (PNG, JPG, BMP via FastOCR), CSV files, OpenXML spreadsheets (XLSX), and plain text files.
- 🔍 Native FastOCR Recognition — Hardware-accelerated image OCR via Windows Media APIs with zero-copy memory management.
- ⚡ Positional Geometry Protection — Uses PDFBox layout extraction with
setSortByPosition(true)and scale-relative visual clustering. - 🚀 Single-Pass RTF Stripper — Fast 0-regex single-pass RTF lexer and control word stripper.
- 📊 StAX Streaming OpenXML — Efficient XML streaming parser for multi-sheet Excel spreadsheets (
.xlsx) supporting shared strings and inline strings. - 🛡️ Binary Guard Protection — Guards against binary
.doc/.docxcorruption with actionable exception feedback.
| Format | Extension | Type | Engine / Strategy | Output |
|---|---|---|---|---|
| Adobe PDF | .pdf |
Document | PDFBox + Visual Paragraph Geometry | Normalized Markdown/Text |
| OpenXML Spreadsheet | .xlsx |
Spreadsheet / Table | StAX Streaming (ZIP + sharedStrings + inlineStr) | TSV / Tabular Text |
| CSV Table | .csv |
Tabular Data | UTF-8 Delimited Line Ingestion & Normalizer | Normalized Text / Grid |
| Rich Text Format | .rtf |
Document | Single-Pass 0-Regex Byte Stripper | Clean Unformatted Text |
| Markdown | .md, .markdown |
Structured Text | Native UTF-8 FastRegex Normalizer | Structured Text |
| Plain Text | .txt, .log |
Unstructured Text | UTF-8 File Reader & Normalizer | Clean Compact Text |
| Image / OCR | .png, .jpg, .bmp |
Visual Media | FastOCR (Hardware Accelerated Windows Media) | Extracted Text |
| MS Word (Legacy) | .doc, .docx |
Binary Word | Guard Protection | Actionable Exception Guidance |
FastContentParse is engineered for high-throughput document ingestion. In the official JMH Benchmark, the system measured raw parsing throughput:
Benchmark Mode Cnt Score Error Units
Benchmark.benchmarkPdfParse thrpt 3 0.248 ± 0.523 ops/ms
Benchmark.benchmarkRtfSinglePassStrip thrpt 3 1274.837 ± 4215.333 ops/ms
1,274,000 Operations per Second: With the single-pass 0-regex RTF stripper,
FastContentParsecleans and normalizes formatted text at over 1.27 Million Operations per Second (1,274 ops/ms). Multi-page PDF text extraction runs with zero memory spikes and scale-relative visual layout clustering.
FastContentParse (This Library — The Parser)
Converts unstructured binary documents (PDF, RTF, Markdown, TXT, XLSX, CSV, OCR images) into normalized UTF-8 text streams.
FastContentChunk (The Strategy Engine)
Segments normalized text streams into contextual passages with Parent-Child context.
FastAIVectorDB (The Vector Store)
High-speed native C++ SIMD vector database storing small chunk.text embeddings for sub-5ms similarity retrieval.
FastAIRag (The Orchestration Pipeline)
Higher-level RAG framework that orchestrates FastContentParse and FastContentChunk, indexes small chunk.text embeddings into FastAIVectorDB, and feeds chunk.parentText to FastAIBot for LLM response generation.
| Method | Description | Path |
|---|---|---|
parseFile(Path) |
Parse a file and auto-detect type by extension. | Reference → |
parseString(String, String) |
Parse raw text and normalize content with inferred type. | Reference → |
parseString(String, String, String) |
Parse raw text with an explicit MIME type. | Reference → |
Run standalone verification demos or execute JMH throughput benchmarks:
| Type | Target / Launcher | Source File | Description |
|---|---|---|---|
| Interactive Demo | run-demo.bat |
Demo.java |
Live PDF/RTF document parsing and visual layout breakdown |
| Throughput Benchmark | run-benchmark.bat |
Benchmark.java |
JMH benchmark evaluating PDF and RTF parsing throughput |
Add the JitPack repository and the dependencies to your pom.xml:
<repositories>
<repository>
<id>jitpack.io</id>
<url>https://jitpack.io</url>
</repository>
</repositories>
<dependencies>
<!-- FastContentParse Core -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastContentParse</artifactId>
<version>0.1.5</version>
</dependency>
<!-- FastJava Ecosystem Dependencies -->
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastRegex</artifactId>
<version>0.1.0</version>
</dependency>
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastOCR</artifactId>
<version>0.1.1</version>
</dependency>
<dependency>
<groupId>com.github.andrestubbe</groupId>
<artifactId>FastCore</artifactId>
<version>0.1.0</version>
</dependency>
<!-- Document Decoding -->
<dependency>
<groupId>org.apache.pdfbox</groupId>
<artifactId>pdfbox</artifactId>
<version>3.0.0</version>
</dependency>
</dependencies>repositories {
mavenCentral()
maven { url 'https://jitpack.io' }
}
dependencies {
// FastContentParse Core
implementation 'com.github.andrestubbe:FastContentParse:0.1.5'
// FastJava Ecosystem Dependencies
implementation 'com.github.andrestubbe:FastRegex:0.1.0'
implementation 'com.github.andrestubbe:FastOCR:0.1.1'
implementation 'com.github.andrestubbe:FastCore:0.1.0'
// Document Decoding
implementation 'org.apache.pdfbox:pdfbox:3.0.0'
}If building manually without Maven or Gradle, include FastContentParse alongside its runtime dependencies on your classpath:
- 📄 FastContentParse-0.1.5.jar — The Core Parser
- ⚡ FastRegex-0.1.0.jar — SIMD Whitespace Normalization
- 👁️ FastOCR-0.1.1.jar — Hardware-Accelerated Image OCR
- ⚙️ FastCore-0.1.0.jar — Unified Native JNI Loader
- 📑 Apache PDFBox 3.0.0+ (
pdfbox-3.0.0.jar,fontbox-3.0.0.jar,commons-logging-1.2.jar) — PDF layout decoding
Important
JitPack (https://jitpack.io) is required to resolve com.github.andrestubbe ecosystem dependencies automatically. When running direct JARs without Maven, ensure all companion JARs above reside on -cp.
- REFERENCE.md: Full API contracts and parser method details.
- PHILOSOPHY.md: Zero-overhead document parsing philosophy.
- COMPILE.md: Maven build instructions.
- CHANGELOG.md: Project history.
- ROADMAP.md: Future development goals.
| Platform | Status |
|---|---|
| Windows 10/11 | ✅ Fully Supported |
| Linux | 🚧 Planned |
| macOS | 🚧 Planned |
MIT License — See LICENSE file for details.
- FastContentChunk — High-performance native SIMD tokenizer and multi-mode strategy chunker
- FastAIVectorDB — High-speed native C++ SIMD vector database
- FastAIRag — Retrieval-Augmented Generation pipeline client
- FastCore — Native JNI loader for FastJava libraries
- FastAI — Unified lightweight AI model client interface
- FastAIModel — Embedded GGUF and ONNX runtimes for local feature embeddings
- FastAIBot — Autonomous conversational AI bot engine
- FastAIAgent — Autonomous agentic workflow execution framework
Part of the FastJava Ecosystem — Making the JVM faster. Small package. Maximum speed. Zero bloat. 🚀📋
