2,066 packages matching “Judge”
@dynatrace-oss/dt-eval-lib
v0.0.17-alpha · 1 day ago
Minimal TypeScript library for running LLM-as-a-judge evaluations
No known vulnerabilities
@dancingteeth/agent-looper
v0.3.0 · 2 days ago
Agent Looper — repo-agnostic fix-until-green harness (Cursor / Cline / OpenCode / Pi / Codex; configurable reviewRuntime judge)
No known vulnerabilities
judge-iphonex
v1.1.0 · 5 years ago
judge is iPhoneX
No known vulnerabilities
@vendoai/guard
v0.36.3 · 3 hours ago
Vendo's policy core at one choke point: guard.bind, approvals, grants, audit, Vendo Auto judge, breakers.
No known vulnerabilities
skill-harness
v0.9.0 · 1 day ago
Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run
No known vulnerabilities
rule-judgment
v1.1.5 · 4 years ago
A query statement similar to mongodb, judge and retrieve data.
No known vulnerabilities
hey-llm-you-okay
v0.2.6 · 1 month ago
Hey LLM, you okay? — a unified, pyramid-ordered LLM testing CLI for CI/CD. YAML-defined layers (static → exec → http → llm → judge), LLM-as-a-judge gates, and an automated failure-triage protocol (A/B probe) that tells you whether a red test is your promp
No known vulnerabilities
@cat-factory/sandbox
v0.12.9 · 3 hours ago
Parallel prompt/model testing surface: versioned prompt candidates, experiment matrices, judge + objective grading. Isolated from the core product so it can be extracted.
No known vulnerabilities
@atbash/atbash-autogen
v0.0.15 · 6 days ago
Atbash safety judge plugin for AutoGen-style multi-agent orchestration
No known vulnerabilities
@nestjs-agentic/evaluation
v1.0.0 · 2 days ago
Automated agent benchmarking, trajectory scoring, and LLM-as-a-Judge evaluation framework for nestjs-agentic.
No known vulnerabilities
@hazeljs/eval
v2.0.3 · 5 hours ago
Evaluation toolkit for HazelJS AI apps — golden datasets, RAG metrics, agent trajectories, LLM-as-judge, CI reports
No known vulnerabilities
opencode-injection-guard
v0.2.1 · 4 months ago
OpenCode plugin that detects prompt injection in tool call outputs using an LLM judge
No known vulnerabilities
pi-until-done
v0.3.1 · 1 month ago
Evidence-driven /until-done goal loops for Pi with TDD planning, mise verification, and mandatory LLM judge gating.
No known vulnerabilities
agentic-test-runner
v0.5.0 · 28 days ago
LLM-judged CLI test runner — define test cases in YAML, execute commands, let an LLM judge PASS/FAIL. Outputs JSONL trace.
No known vulnerabilities
@looprun-ai/eval
v0.20.0 · 10 days ago
looprun eval harness: run a generated subject's cases against a model target (governed or ungoverned variant), dump per-case traces for the LLM judge, fold verdicts, and certify against the bar.
No known vulnerabilities
@reactive-agents/eval
v0.15.0 · 4 days ago
Evaluation framework for Reactive Agents — LLM-as-judge scoring, regression detection, dataset loading
No known vulnerabilities
acyclic-eval
v0.1.4 · 20 days ago
Evaluate LLM and rule-based judges with mutation cases whose generation does not depend on the judge under test.
No known vulnerabilities
@alexeiled/pi-fusion
v0.8.0 · 6 hours ago
Stronger answers for hard Pi questions via a parallel model panel + judge, built on pi-subagents
No known vulnerabilities
judge-d
v1.5.1 · 4 years ago
CLI for publishing and validating contract tests using judge-d API
No known vulnerabilities
@eva-llm/eva-judge
v1.0.8 · 1 month ago
LLM-as-a-Judge abstraction layer using ai-sdk and plugins
No known vulnerabilities