477,083 packages matching “judge-d”
@memberjunction/computer-use
v5.51.1 · 2 days ago
MemberJunction: Computer Use - Vision-to-Action engine for driving web browsers via LLM reasoning over screenshots. MJ-independent base layer.
No known vulnerabilities
nave
v3.5.6 · 7 months ago
Virtual Environments for Node
No known vulnerabilities
language-loop
v0.5.0 · 4 days ago
A localization loop for vibe-coding agents. Scans code, extracts i18n keys, translates only what changed, and automatically holds unsafe translations back.
No known vulnerabilities
@handsealed/verifier
v0.25.0 · 13 days ago
The Handsealed CLI: replay every verdict offline against a clone. Don't trust us — verify.
No known vulnerabilities
acyclic-eval
v0.1.4 · 20 days ago
Evaluate LLM and rule-based judges with mutation cases whose generation does not depend on the judge under test.
No known vulnerabilities
@a3s-lab/sentry
v0.3.0 · 17 days ago
Native (in-process) SDK for a3s-sentry — judge observer events through the embedded L1/L2/L3 pipeline.
No known vulnerabilities
skill-harness
v0.9.0 · 1 day ago
Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run
No known vulnerabilities
@dancingteeth/agent-looper
v0.3.0 · 2 days ago
Agent Looper — repo-agnostic fix-until-green harness (Cursor / Cline / OpenCode / Pi / Codex; configurable reviewRuntime judge)
No known vulnerabilities
agentic-sage
v1.3.1 · 24 days ago
Passive fleet judge for parallel AI coding agent sessions (Claude Code, Grok Build CLI, etc.) — board, territory, merge briefings, optional live-judge briefs.
No known vulnerabilities
mongodb-assistant-eval
v0.0.8 · 5 months ago
Evaluation library for the MongoDB Assistant API.
No known vulnerabilities
eval-bench
v0.25.0 · 24 days ago
Benchmark Claude Code plugins by A/B comparing plugin versions with LLM-judged evaluation prompts.
No known vulnerabilities
deepeval
v0.9.11 · 1 day ago
The LLM Evaluation Framework for TypeScript
No known vulnerabilities
@memberjunction/computer-use-engine
v5.51.1 · 2 days ago
MemberJunction: MJ Computer Use - MJ-aware subclass of the Computer Use engine with AIPromptRunner, MJ Credentials, Actions-as-Tools, and entity persistence.
No known vulnerabilities
@quarkos/pi-fusion
v1.0.8 · 1 month ago
Multi-model deliberation harness (OpenRouter Fusion pattern) for OpenCode Go
No known vulnerabilities
react-scrollbars-custom
v4.1.1 · 4 years ago
The best React custom scrollbars component
No known vulnerabilities
@mzwing/pi-permission-auto-review
v0.2.0 · 11 days ago
Codex-style automatic approval reviews for @gotgenes/pi-permission-system
No known vulnerabilities
@skill-harness/cli
v0.9.0 · 1 day ago
skill-harness CLI — run, grade, review, and lint agent-skill scenarios on the pi harness
No known vulnerabilities
@remnic/bench
v9.69.14 · 7 minutes ago
Retrieval latency ladder benchmarks + CI regression gates for @remnic/core
No known vulnerabilities
@gotgenes/pi-permission-model-judge
v1.1.4 · 2 days ago
Deny-first typo-path model judge — a pi-permission-system Authorizer chain link
No known vulnerabilities
@dynatrace-oss/dt-evals
v0.2.22-alpha · 1 day ago
Evaluation CLI for AI Observability on Dynatrace
No known vulnerabilities