33,085 packages matching “llm-judge”
llm-judge
v1.0.0 · 4 months ago
CLI tool for evaluating artifacts with LLMs — Swiss Elo ranking, pass/gate, and qualitative review
No known vulnerabilities
llm-scorer
v0.1.3 · 1 month ago
A tiny JSONL LLM judge
No known vulnerabilities
@skill-harness/core
v0.17.0 · 5 days ago
skill-harness engine — spec, discover, run, LLM-judge grade, score, results (internal API)
No known vulnerabilities
@spendgraph/evals
v0.8.3 · 1 day ago
Score LLM output instead of asserting equality. Deterministic metrics, an optional LLM judge, and the bill for both.
No known vulnerabilities
skill-harness
v0.17.0 · 5 days ago
Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run
No known vulnerabilities
opencode-injection-guard
v0.2.1 · 5 months ago
OpenCode plugin that detects prompt injection in tool call outputs using an LLM judge
No known vulnerabilities
llm-governance-gateway
v0.15.0 · 5 days ago
Governed structured-output LLM pipeline: rate limit → spend caps (per-user + global circuit breaker) → cache → provider-chain failover → Zod validation → usage logging → LLM-judge, with a deterministic mock provider for CI.
No known vulnerabilities
dsh-write-gate
v0.1.1 · 1 month ago
Commitment write-gate for AI coding agents: two-tier (deterministic + LLM judge) pre-execution policy. Engine-agnostic core with a DeepSeek Harness (dsh) adapter.
No known vulnerabilities
@pwtap/plugin-ai-judge
v0.2.0 · 1 month ago
AI/LLM judge matchers for Playwright — toPassRubric/toScoreAtLeast/toMatchImage over Ollama, OpenAI-compatible endpoints (OpenAI/OpenRouter/NVIDIA/Groq/…), and native Claude
No known vulnerabilities
@unotest/judge
v0.40.0 · 1 day ago
LLM-judge service for the @unotest ecosystem: judges free-form text (chat replies, generated content) against a natural-language rubric and returns a structured pass/fail verdict with reasoning. Runs as a small HTTP service (`npx @unotest/judge`) or in-pr
No known vulnerabilities
@reaatech/llm-judge-cache
v0.1.1 · 3 months ago
Multi-backend caching for LLM Judge Toolkit
No known vulnerabilities
hey-llm-you-okay
v0.2.6 · 2 months ago
Hey LLM, you okay? — a unified, pyramid-ordered LLM testing CLI for CI/CD. YAML-defined layers (static → exec → http → llm → judge), LLM-as-a-judge gates, and an automated failure-triage protocol (A/B probe) that tells you whether a red test is your promp
No known vulnerabilities
@adia-ai/gen-ui
v0.8.66 · 6 days ago
AdiaUI generative-UI system — compose strategies, retrieval, the harvested training corpus, and catalog-aware + LLM-judge validation. Emits A2UI protocol messages; pairs with @adia-ai/a2ui (the protocol runtime). Folded from @adia-ai/a2ui-{compose,retriev
No known vulnerabilities
@reaatech/llm-judge-templates
v0.1.1 · 3 months ago
Prompt templates for LLM Judge Toolkit evaluation criteria
No known vulnerabilities
pi-until-done
v0.3.1 · 2 months ago
Evidence-driven /until-done goal loops for Pi with TDD planning, mise verification, and mandatory LLM judge gating.
No known vulnerabilities
@reaatech/llm-judge-consensus
v0.1.1 · 3 months ago
Multi-judge consensus strategies for LLM Judge Toolkit
No known vulnerabilities
pi-shift-router
v1.6.0 · 2 days ago
An LLM judge routes every Pi agent turn to the right model — fast execution for routine work, smart reasoning for hard problems — with multi-model failover and automatic orchestration.
No known vulnerabilities
@tagma/completion-llm-judge
v0.2.96 · 1 hour ago
LLM-as-judge completion plugin for tagma-sdk pipelines
No known vulnerabilities
@reaatech/llm-judge-types
v0.1.1 · 3 months ago
Core types, Zod schemas, and errors for LLM Judge Toolkit
No known vulnerabilities
@reaatech/llm-judge-infra
v0.1.1 · 3 months ago
Cost tracking, monitoring, and batch processing for LLM Judge Toolkit
No known vulnerabilities