2,073 packages matching “Judge”
code-block-judge
v1.0.6 · 2 years ago
Judge the code block information in the markdown
No known vulnerabilities
@judge-java/cli
v0.5.2 · 4 months ago
CLI for judge-java — init, dashboard, and management commands
No known vulnerabilities
@artale/pi-eval
v1.3.2 · 4 months ago
Agent evaluation harness. Judge sessions on success, tool usage, efficiency, methodology. Inspired by opencc.
No known vulnerabilities
is-bool
v0.0.3 · 9 years ago
Judge the condition
No known vulnerabilities
project-loop
v2.0.0 · 24 days ago
Closed-loop, evidence-gated build system for AI coding agents. Plan, Spec, Build, Verify — with a separate Judge role that holds the exit gate and issues rework orders until the frozen Definition of Done is met.
No known vulnerabilities
jstypecgf
v1.1.0 · 6 years ago
javascript type judge
No known vulnerabilities
@runflow-ai/evals
v0.0.9 · 1 month ago
Runflow Evals — project-local evals framework for Runflow agents (datasets, scorers, journey/conversation validation, LLM judge, viewer)
No known vulnerabilities
llmt-runner
v0.0.1 · 29 days ago
LLM-judged CLI test runner — define test cases in YAML, execute commands, let an LLM judge PASS/FAIL. Outputs JSONL trace.
No known vulnerabilities
pulse-flight-deals-mcp
v0.3.2 · 11 days ago
MCP server that lets an AI agent judge whether a flight price is actually a deal, grounded in observed fare history. Answers 'is $X for this route a good price?' with real numbers an LLM cannot produce on its own.
No known vulnerabilities
@atbash/mcp
v0.1.3 · 3 months ago
Atbash safety judge exposed as a standalone MCP server
No known vulnerabilities
@they-juanreina/compost-evals
v0.2.1 · 2 months ago
Eval surfaces: versioned LLM-as-judge rubric, eval-grader loop, skill golden-set runner.
No known vulnerabilities
@reaatech/llm-judge-templates
v0.1.1 · 2 months ago
Prompt templates for LLM Judge Toolkit evaluation criteria
No known vulnerabilities
@pharmatools/opengate-mcp
v0.1.4 · 1 month ago
MCP server for OpenGATE — let an AI agent check whether its answers are grounded in the provided context. Deterministic, no LLM judge.
No known vulnerabilities
openfusion-mcp
v0.3.1 · 1 month ago
Local MCP server implementing OpenRouter's Fusion panel architecture: fan-out + two-step judge.
No known vulnerabilities
@swoosh-dev/judge
v0.2.0 · 2 months ago
Dynamic routing policies for @swoosh-dev/router — classify the prompt with an LLM judge (structured output) and route accordingly.
No known vulnerabilities
@taskproof/judge
v0.2.1 · 2 months ago
Optional LLM judge (WebJudge-style) for taskproof: a versioned rubric prompt that grades goal completion from a run's evidence, gated behind the deterministic assertions
No known vulnerabilities
bojhelper
v2.0.1 · 6 years ago
A tool for Baekjoon Online Judge
No known vulnerabilities
@varlabs/ai.evals
v0.1.0 · 12 days ago
Dataset regression evals for the AI SDK: pluggable scorers (exact-match, embedding similarity, LLM-as-judge) that compose with your own vitest suite
No known vulnerabilities
judge-solve0
v78.3.880 · 2 years ago
judge-solve0
No known vulnerabilities
@reaatech/llm-judge-calibration
v0.1.1 · 2 months ago
Calibration metrics, runner, datasets, and drift detection for LLM Judge Toolkit
No known vulnerabilities