361,087 packages matching “text extraction”
rebulk-js
v3.4.1 · 2 days ago
A generic pattern matching engine for rule-based text extraction. TypeScript port of Python rebulk.
No known vulnerabilities
@m5kdev/module-pdf
v0.37.4 · 14 days ago
Optional Backend Module for PDF text extraction.
No known vulnerabilities
@hazeljs/pdf-to-audio
v2.0.7 · 16 days ago
Convert PDF documents to audio using TTS - PDF text extraction, chunking, and OpenAI TTS
No known vulnerabilities
dao-zang-skill
v2.0.0 · 28 days ago
Installable bundle contributing the DaoZang offline retrieval & original-text extraction skill to DeepSeek Harness
No known vulnerabilities
pptx-browser
v4.1.5 · 6 months ago
Render, edit, and export PowerPoint (PPTX) slides. Canvas rendering, SVG export, PDF export, PPTX writer/template engine, slide show, text extraction, animations. Zero dependencies.
No known vulnerabilities
pdf-oxide
v0.3.77 · 1 month ago
The fastest Node.js PDF library — 0.8ms mean, 5× faster than the industry leaders, 100% pass rate on 3,830 real-world PDFs. Prebuilt native bindings, no build toolchain: text extraction, Markdown/HTML conversion, PDF creation and editing.
No known vulnerabilities
@fradser/pi-kit
v0.5.1 · 2 days ago
Shared runtime helpers for FradSer pi packages — TUI spinner/theme primitives and message text extraction. Internal workspace dependency, not a Pi package.
No known vulnerabilities
@remit/attachment-service
v0.0.8 · 1 month ago
Pure text extraction from email attachment bytes (PDF, DOCX, DOC, plain text)
No known vulnerabilities
n8n-nodes-tesseractjs7
v2.6.4 · 13 days ago
An n8n community node package for memory-efficient sequential PDF text extraction, OCR preflight, recognition, and lightweight page slicing
No known vulnerabilities
pdf-oxide-wasm
v0.3.77 · 1 month ago
The fastest WebAssembly PDF library — 0.8ms mean, 5× faster than the industry leaders, 100% pass rate on 3,830 real-world PDFs. Zero-dependency for Node.js, browsers, and edge runtimes: text extraction, Markdown/HTML conversion, search, form filling, crea
No known vulnerabilities
@ismail-elkorchi/html-parser
v0.2.1 · 1 month ago
HTML parser with bounded text extraction, fragment parsing, and structural traversal.
No known vulnerabilities
@fabriqa.ai/pdf-reader-mcp
v1.0.7 · 10 months ago
MCP server for efficient PDF text extraction, search, and metadata retrieval for Claude Code
No known vulnerabilities
exvoluptate
v1.6.0 · 2 years ago
PDF text extraction in TypeScript
No known vulnerabilities
@ohos-ports/llamaindex-liteparse
v2.5.1-beta.2 · 1 month ago
Fast, lightweight PDF and document parsing with spatial text extraction
No known vulnerabilities
@dvvebond/core
v0.2.17 · 6 months ago
A fork of @libpdf/core with enhanced React components, Azure Document Intelligence integration, text extraction with bounding boxes, and enterprise PDF viewing features
No known vulnerabilities
node-extract-text-from-file
v2.0.2 · 5 years ago
Detection and text extraction supported for .DOC, .DOCX, .PDF files
No known vulnerabilities
@intelagent/mcp-file-processor
v0.1.1 · 6 months ago
MCP server for universal text extraction, keyword extraction, language detection, and text chunking for RAG pipelines
No known vulnerabilities
documelt
v0.2.0 · 1 month ago
WASM-powered document text extraction for PDF, DOCX, XLSX, PPTX
No known vulnerabilities
@zzwz/liteparse-vllm
v1.5.3-custom.1 · 4 months ago
Open-source PDF parsing with spatial text extraction and OCR processing with Custom Codex-OCR and GML-OCR Servers
No known vulnerabilities
tika
v1.6.1 · 9 years ago
Apache Tika bridge. Text extraction, metadata extraction, mimetype detection and language detection.
No known vulnerabilities