132,122 packages matching “ai-judge”
ai-judge
v0.0.1 · 10 months ago
Whitecircle.ai tools for judge-style evaluations and rubric scoring.
No known vulnerabilities
@caslanqa/create-playwright-ai
v4.0.1 · 25 days ago
Scaffold a ready-to-run Playwright + TypeScript test framework: AI Judge for LLM evaluation, layered API testing, lazy session auth, and full tooling (ESLint, Prettier, husky, commitlint)
No known vulnerabilities
@pwtap/plugin-ai-judge
v0.1.2 · 17 days ago
AI/LLM judge matchers for Playwright — toPassRubric/toScoreAtLeast/toMatchImage over Ollama, OpenAI-compatible endpoints (OpenAI/OpenRouter/NVIDIA/Groq/…), and native Claude
No known vulnerabilities
chance-mcp
v1.0.0 · 8 days ago
Chance verification harness as an MCP tool — the AI judge between an agent's intent and its action. Verify a proposed action against its mandate before it executes, and get a signed, independently-verifiable verdict.
No known vulnerabilities
vitest-evals
v0.16.1 · 4 days ago
Harness-backed AI testing on top of Vitest.
No known vulnerabilities
deepthink-js
v1.4.0 · 1 month ago
SOTA NPM module for agentic processes using local or cloud LLMs.
No known vulnerabilities
@metaharness/redblue
v0.1.4 · 1 month ago
AI red-teaming for the AI agents & LLM apps you own: stress-test them with adversarial models to find security failures (prompt injection, tool misuse / excessive agency, data leakage, jailbreaks, denial-of-wallet), auto-patch (blue team), retest, and get
No known vulnerabilities
openevals
v0.2.0 · 4 months ago
Much like tests in traditional software, evals are an important part of bringing LLM applications to production. The goal of this package is to help provide a starting point for you to write evals for your LLM applications, from which you can write more c
No known vulnerabilities
idea-judge
v1.2.0 · 4 months ago
Spawn 10 AI judge personas to evaluate competing ideas and pick the best one
No known vulnerabilities
agent-testing-library
v1.5.1 · 8 months ago
Testing framework for AI agents and multi-agent systems with AI Judge verification (Node & Bun compatible)
No known vulnerabilities
@looprun-ai/eval
v0.20.0 · 1 day ago
looprun eval harness: run a generated subject's cases against a model target (governed or ungoverned variant), dump per-case traces for the LLM judge, fold verdicts, and certify against the bar.
No known vulnerabilities
mongodb-assistant-eval
v0.0.8 · 5 months ago
Evaluation library for the MongoDB Assistant API.
No known vulnerabilities
playwright-ai-triage
v0.9.0 · 1 day ago
Playwright reporter that classifies test failures with an LLM (real bug / flaky / selector drift / environment issue) and posts a readable summary to stdout, a GitHub PR comment, or Slack.
No known vulnerabilities
@framers/agentos
v0.10.14 · 4 days ago
AgentOS: open-source TypeScript framework for autonomous AI agents. Unified graph orchestration, cognitive memory, runtime tool forging, multi-tier guardrails, voice pipeline, and 11 LLM providers.
No known vulnerabilities
nanoid
v6.0.1 · 8 days ago
A tiny (118 bytes), secure URL-friendly unique string ID generator
No known vulnerabilities
@presentation-md/render
v1.20.9 · 7 days ago
Render a deck JSON spec to a self-contained HTML slide deck.
No known vulnerabilities
@ai-sdk/gateway
v4.0.49 · 7 hours ago
The Gateway provider for the [AI SDK](https://ai-sdk.dev/docs) allows the use of a wide variety of AI models and providers.
No known vulnerabilities
deepeval
v0.9.10 · 2 days ago
The LLM Evaluation Framework for TypeScript
No known vulnerabilities
@introspection-ai/cli
v0.25.2 · 6 hours ago
CLI for operating agents on the Introspection platform
No known vulnerabilities
@tracecode/harness
v0.16.4 · 2 hours ago
Browser-native execution for Python, JavaScript, TypeScript, Java, C#, and C++.
No known vulnerabilities