477,084 packages matching “judge-d”
convex-evalbench
v0.3.0 · 1 month ago
Reactive LLM eval, tracing, and regression layer as a Convex component.
No known vulnerabilities
jupyterlab-judge
v1.31.0 · 7 days ago
A simple online judge for Jupyter Lab.
No known vulnerabilities
pi-until-done
v0.3.1 · 1 month ago
Evidence-driven /until-done goal loops for Pi with TDD planning, mise verification, and mandatory LLM judge gating.
No known vulnerabilities
@gaunt-sloth/batch
v2.0.0-beta.0 · 2 days ago
Batch matrix runtime for Gaunt Sloth (BATCH-1): run a prompt-executable over a matrix of models and/or content-bound inputs
No known vulnerabilities
d
v1.0.2 · 2 years ago
Property descriptor factory
No known vulnerabilities
cuberyl
v1.0.0 · 5 years ago
NxNxN cube puzzle simulator
No known vulnerabilities
@homepages/template-kit
v7.0.0 · 5 days ago
Authoring kit for HomePages marketing-section templates: schema system, contract primitives, theme tokens, and the design-system CSS layer.
No known vulnerabilities
postcss-viewport-units
v0.1.6 · 7 years ago
Automatically append `content` property for viewport-units-buggyfill
No known vulnerabilities
@mizore66/faultline
v0.1.2 · 1 month ago
Executable evidence for regressions introduced during Codex-assisted development.
No known vulnerabilities
@ust-protocol/cli
v1.0.0-rc.106 · 5 days ago
The reference `ust` CLI for UST (Universal State Transcript) 1.0 — verify a transcript, print canonical bytes for cross-language diffing, and run the HIGH genesis ceremony. One entrypoint; the Go binary reproduces this surface.
No known vulnerabilities
playwright-ai-triage
v0.9.1 · 4 days ago
Playwright reporter that classifies test failures with an LLM (real bug / flaky / selector drift / environment issue) and posts a readable summary to stdout, a GitHub PR comment, or Slack.
No known vulnerabilities
@maestrofrontier/frontier
v1.15.0 · 25 days ago
Maestro Frontier is an opt-in, zero-dependency local multi-CLI fusion engine for AI coding agents. Compose an Opus 5 and Codex model panel, set provider-supported effort, judge its one-shot read-only responses, and synthesize one grounded answer. Optional
No known vulnerabilities
@introspection-ai/cli
v0.30.0 · 1 day ago
CLI for operating agents on the Introspection platform
No known vulnerabilities
pi-fusion
v0.9.0 · 26 days ago
Multi-model deliberation for pi, inspired by OpenRouter Fusion
No known vulnerabilities
opencode-tamer-judge
v1.1.0 · 20 days ago
LLM-powered judge for opencode - blocks file modifications without explicit user permission
No known vulnerabilities
@fusionkit/cli
v0.10.1 · 26 days ago
fusionkit — real model fusion behind your coding agent (Codex, Claude Code, Cursor, OpenCode).
No known vulnerabilities
tinyrainbow
v3.1.1 · 24 days ago
A small library to print colourful messages.
No known vulnerabilities
@webappwiz/cli
v0.0.3 · 2 days ago
The webappwiz CLI: judge code against rules, sign off a diff, and manage agent skills
No known vulnerabilities
@lythos/test-utils
v0.17.3 · 21 days ago
    mode permission classifier for DeepSeek Harness: a Claude-Code-auto-mode-like classifier over tools/pre-execute and approval/request, a selectable 'auto' permission preset, LLM semantic judge, git checkpointing, agent discipline guidance
No known vulnerabilities