240,307 packages matching “html-extraction”
@markdownee/trafilaturacore
v0.8.2 · 2 days ago
Offline supplied-HTML extraction, cleaning, and sanitization in pure TypeScript.
No known vulnerabilities
@advance-labs/html-parser
v0.2.2 · 13 days ago
Pure HTML extraction for the AEO Toolkit — meta/OG/Twitter, headings, images, links, content signals, and raw structured-data blocks. No network.
No known vulnerabilities
911-scraper-mcp
v1.0.9 · 9 months ago
911Proxy Universal Web Scraper MCP Server - supports HTML extraction and screenshots
No known vulnerabilities
rewriterkit
v1.0.2 · 1 month ago
Declarative HTML extraction library built on Cloudflare HTMLRewriter.
No known vulnerabilities
tar
v7.5.22 · 2 months ago
tar for node
No known vulnerabilities
@zhengxs/temme
v0.1.1 · 24 days ago
HTML extraction DSL rewritten on a POJO execution plan with a linkedom default adapter.
No known vulnerabilities
pi-local-rag
v0.4.1 · 3 months ago
Local hybrid RAG pipeline for the Pi coding agent. SQLite FTS5 + sqlite-vec, ONNX embeddings via Transformers.js, PDF/DOCX/HTML extraction (with OCR fallback), per-project storage, auto-injection. Zero cloud dependency.
No known vulnerabilities
@haybarn/ext-crawler-h1-5-3
v202602.17.1216463 · 3 months ago
SQL-native web crawler with HTML extraction and MERGE support
No known vulnerabilities
webtester-mcp
v0.2.0 · 10 months ago
MCP server for website testing with Playwright - browser automation, HTML extraction, JavaScript execution, monitoring
No known vulnerabilities
@haybarn/ext-crawler-h1-5-4
v202602.17.121646 · 3 months ago
SQL-native web crawler with HTML extraction and MERGE support
No known vulnerabilities
@haybarn/ext-crawler-h1-5-2
v202605.18.140429 · 4 months ago
SQL-native web crawler with HTML extraction and MERGE support
No known vulnerabilities
pip-requirements-js
v1.0.3 · 6 months ago
A robust parser for requirements.txt files
No known vulnerabilities
deeks
v3.2.1 · 4 months ago
Retrieve all keys and nested keys from objects and arrays of objects.
No known vulnerabilities
@nosferatu500/textract
v3.1.3 · 3 years ago
Extracting text from files of various type including html, pdf, doc, docx, xls, xlsx, csv, pptx, png, jpg, gif, rtf, text/*, and various open office.
No known vulnerabilities
@lingui/babel-plugin-extract-messages
v6.8.0 · 14 hours ago
Babel plugin to extract translatable messages from source code into Lingui catalogs
No known vulnerabilities
apparatus
v0.0.10 · 8 years ago
various machine learning routines for node
No known vulnerabilities
pi-web-access
v0.31.0 · 22 hours ago
Web search, URL fetching, GitHub repo cloning, PDF extraction, YouTube video understanding, and local video analysis for Pi coding agent. Supports OpenAI, Brave, Parallel, TinyFish, Search1API, Searchinfinity, Querit, Tavily, Firecrawl, Crawl4AI, Jina, SE
No known vulnerabilities
textract
v2.5.0 · 7 years ago
Extracting text from files of various type including html, pdf, doc, docx, xls, xlsx, csv, pptx, png, jpg, gif, rtf, text/*, and various open office.
1 critical
strong-globalize
v6.0.6 · 3 years ago
StrongLoop Globalize - API
No known vulnerabilities
@wordpress/dependency-extraction-webpack-plugin
v6.56.0 · 16 hours ago
Extract WordPress script dependencies from webpack bundles.
No known vulnerabilities