589 packages matching “judges”
@kevinrabun/judges
v3.129.9 · 5 months ago
45 specialized judges that evaluate AI-generated code for security, cost, and quality.
No known vulnerabilities
judges
v0.0.1 · 10 months ago
Coming soon.
No known vulnerabilities
@kevinrabun/judges-cli
v3.129.9 · 5 months ago
CLI wrapper for the Judges code review toolkit.
No known vulnerabilities
@post-print/agent-test
v2.0.6 · 6 days ago
Playwright-based coding-agent tests, comparisons, explicit judges, and viewer
No known vulnerabilities
@tangle-network/agent-eval
v0.187.2 · 12 hours ago
Evaluate and improve AI agents from runs, traces, judges, and feedback. Compare candidates, cluster failures, measure lift, and gate releases.
No known vulnerabilities
backant-kairos
v1.6.0 · 10 days ago
Autonomous AI engineer — observes, judges, builds, ships
No known vulnerabilities
ohbem
v1.5.3 · 2 years ago
Ohbem judges your Pokemon GO IVs.
No known vulnerabilities
@matatbread/matbot-triggers
v0.4.16 · 6 days ago
Data-driven hooks: stored conditions that, when an LLM classifier judges them matched, invoke a tool. Cross-runtime (node + browser).
No known vulnerabilities
@prismatic-io/lux
v0.0.1 · 1 month ago
Coding-agent evaluation and improvement with deterministic assertions, optional LLM judges, and first-class human-in-the-loop runs.
No known vulnerabilities
@gethmy/harness
v1.19.2 · 2 hours ago
Execution motor for Harmony playbook stages. Runs exactly one stage per invocation: worktree, role-separated subagents, held oracle, gate evidence. It never routes, never judges, never pushes.
No known vulnerabilities
@nannier-com/lookout
v0.59.6 · 15 days ago
Project-agnostic visual AI tester: captures what an app actually renders, judges it against UI best practices and per-project rules via the local Claude Code CLI, verifies acceptance criteria, and tracks findings in an adjudicated backlog.
No known vulnerabilities
@nullius-inverba/kit
v0.7.0 · 18 days ago
Witness recording for agent runs — harness hooks emit a journal the agent cannot decline to write, and `nullius witness validate` judges it.
No known vulnerabilities
pi-warden
v0.48.1 · 14 hours ago
Makes the Pi agent follow your project's rules. Jev judges every write against your pi-warden.md and quotes the broken rule back to the agent, names slop, breaks stuck loops, calls out unverified done claims, compresses large tool output, and holds the ra
No known vulnerabilities
@unotest/judge
v0.42.0 · 1 hour ago
LLM-judge service for the @unotest ecosystem: judges free-form text (chat replies, generated content) against a natural-language rubric and returns a structured pass/fail verdict with reasoning. Runs as a small HTTP service (`npx @unotest/judge`) or in-pr
No known vulnerabilities
@ggui-ai/negotiator
v0.22.0 · 1 day ago
Contract-synthesis + match-judge engine for ggui's handshake. Synthesizes or repairs a conforming DataContract from an agent's draft, judges blueprint-match candidates for reuse, and validates contract structure + novelty — the primitives composed by deci
No known vulnerabilities
@clossys/inspector
v0.2.8 · 1 day ago
Judges a change before it lands: a secret-scan attempt (attested, or actually run via the ./secret-scan subpath's verified gitleaks acquisition), a change's task record, its review evidence, and policy drift, each reported as satisfied/violated/indetermin
No known vulnerabilities
@agentv/sdk
v4.42.4 · 3 months ago
Evaluation SDK for AgentV - build custom code judges
No known vulnerabilities
compact-adviser
v0.1.8 · 9 hours ago
Tells you when your Pi session has reached a good point to compact: a standalone extension that judges completed checkpoints instead of leaving you to guess.
No known vulnerabilities
guard-my-design-system
v1.9.0 · 4 hours ago
Your design system dies one pull request at a time. This makes sure it doesn't. A guard that judges only the lines a change adds, against the system the repo already has, and names the on-system value the author probably meant.
No known vulnerabilities
@max-null/dsh-habit
v0.2.0 · 1 month ago
Self-learning habit engine for the DeepSeek Harness — detects user-correction signals, judges habits with a low-cost model on threshold, settles candidates behind a two-level human gate
No known vulnerabilities