32 packages matching “unigram”
unigram
v1.1.1 · 5 years ago
The 1/3 million most frequent words, all lowercase, with counts.
No known vulnerabilities
kitoken
v0.11.0 · 4 months ago
Fast tokenizer for language models, supporting BPE, Unigram and WordPiece tokenization
No known vulnerabilities
@verbx/localize-offline
v0.2.2 · 27 days ago
VerbX Localize — offline ONNX translation: Unigram tokenizer + encoder/decoder generation, no network.
No known vulnerabilities
@flexpilot-ai/tokenizers
v0.0.1 · 1 year ago
Node.js binding for huggingface/tokenizers library
No known vulnerabilities
t9-plus
v3.1.3 · 5 years ago
Word prediction for T9 keyboard.
No known vulnerabilities
@mailwoman/neural-weights-en-us
v10.0.0 · 14 days ago
Mailwoman neural-classifier weights for locale 'en-us'. Data-only package — loaded by @mailwoman/neural at runtime.
No known vulnerabilities
languagemodel
v0.4.0 · 12 years ago
A natural language model and cross-language model, for natural language understanding and generation
No known vulnerabilities
@mailwoman/neural
v10.0.0 · 14 days ago
Mailwoman neural classifier runtime: SentencePiece tokenizer + ONNX inference + decoder wiring.
No known vulnerabilities
three-llm
v0.5.2 · 8 days ago
Browser LLM inference on WebGPU using Three.js TSL compute shaders.
No known vulnerabilities
pi-ruleset
v0.3.1 · 1 month ago
Store and apply business rules as structured per-day Markdown files with BM25 search
No known vulnerabilities
@dsh-enhanced/personal-wiki
v0.1.48 · 22 hours ago
A safe, searchable, approval-gated Markdown knowledge vault for DeepSeek Harness.
No known vulnerabilities
@mailwoman/neural-weights-fr-fr
v10.0.0 · 14 days ago
Mailwoman neural-classifier weights for locale 'fr-fr'. Data-only overlay — shares the base model.onnx + tokenizer.model from @mailwoman/neural-weights-en-us (they are byte-identical); ships only the fr-specific data siblings. Loaded by @mailwoman/neural
No known vulnerabilities
@dsh-enhanced/personal-memory
v0.1.48 · 22 hours ago
Scoped, bounded, approval-gated long-term memory for DeepSeek Harness agents.
No known vulnerabilities
tokengeex
v0.6.2 · 2 years ago
This repository holds the code for the TokenGeeX Rust crate and Python package. TokenGeeX is a tokenizer for [CodeGeeX](https://github.com/THUDM/Codegeex2) aimed at code and Chinese. It is based on [UnigramLM (Taku Kudo 2018)](https://arxiv.org/abs/1804.1
No known vulnerabilities
create-llm
v2.2.3 · 11 months ago
The fastest way to start training your own Language Model. Create production-ready LLM training projects in seconds.
No known vulnerabilities
@theanikrtgiri/create-llm
v2.2.3 · 11 months ago
The fastest way to start training your own Language Model. Create production-ready LLM training projects in seconds.
No known vulnerabilities
@khoralabs/tkn
v0.1.0 · 1 month ago
Fast greedy pattern discovery and lattice-backed tokenizer for sequential data.
No known vulnerabilities
word-ngrams
v0.2.0 · 11 years ago
A package for building and analyzing word nGrams
No known vulnerabilities
kr-history-bm25
v0.5.0 · 2 months ago
Convert Korean historical source XML (국사편찬위원회 한국사DB) into a searchable SQLite BM25 corpus — hanja-primary full-text search, structured place/person index, co-occurrence clustering, and an optional LLM literal-translation secondary index.
No known vulnerabilities
ai-token-estimator
v1.7.1 · 8 months ago
Estimate and count tokens (incl. exact OpenAI BPE) and input costs for LLM API calls
No known vulnerabilities