58 packages matching “promptfoo”
openeval-sdk
v1.0.0 · 8 days ago
TypeScript SDK for OpenEval
No known vulnerabilities
@eva-llm/eva-judge
v1.0.8 · 1 month ago
LLM-as-a-Judge abstraction layer using ai-sdk and plugins
No known vulnerabilities
whatbroke-cli
v0.5.0 · 12 days ago
Diff your AI agent's behavior between two runs. Swap a model, change a prompt, run whatbroke, see exactly what changed.
No known vulnerabilities
zeroleaks
v1.4.0 · 29 days ago
AI Security Scanner - Test your AI systems for prompt injection and extraction vulnerabilities
No known vulnerabilities
agent-skill-harbor-collector
v0.15.6 · 4 months ago
Collector and post-collect runtime package for Agent Skill Harbor.
No known vulnerabilities
@eva-llm/eva-cli
v1.0.7 · 3 months ago
A terminal-based tool for local runs and debugging of eva-run
No known vulnerabilities
phi-leak-guard
v1.0.0 · 12 days ago
Deterministic PHI-leak assertions for LLM output, runnable in your normal test suite. HIPAA Safe Harbor + NHS/UK. No LLM calls, no data leaves your machine.
No known vulnerabilities
prompt-iteration-assistant
v0.0.37 · 2 years ago
A set of CLI tools to help you iterate on your LLM prompts.
No known vulnerabilities
promptfoo_sql
v0.17.2 · 3 years ago
LLM eval & testing toolkit
No known vulnerabilities
@wartzar-bee/promptdrift
v0.1.0 · 2 months ago
Catch prompt regressions from model drift — on a schedule, not just on PRs. Runs a small eval set against your LLM on a cron, compares pass-rate to a stored baseline, and alerts (GitHub issue + non-zero exit) when it regresses.
No known vulnerabilities
mapterrain
v0.4.0 · 16 days ago
Terrain test intelligence CLI.
No known vulnerabilities
@iflow-mcp/promptfoo-evil-mcp-server
v1.0.4 · 5 months ago
MCP server that simulates malicious behaviors for security testing
No known vulnerabilities
@hyuga/genchi
v0.2.0 · 11 days ago
Completion verification gate for AI agents & automation: before anything reports "done", re-fetch the real world state with a probe and block the completion claim on empty/mismatch/error. Framework-agnostic, zero-dependency, no LLM.
No known vulnerabilities
evaldrift
v0.1.0 · 1 month ago
Snapshot-style regression testing for LLM prompts & agents — catch silent quality drift when you tweak a prompt. China-model & Chinese-first (DeepSeek / Kimi / Qwen / Doubao / Ollama).
No known vulnerabilities
@idriszade/eval
v0.1.9 · 2 months ago
Pipeline-kit eval foundation — defineEval, runEval, case/scorer/score types
No known vulnerabilities
@jvrmaia/chatlab
v0.3.0 · 2 months ago
Local development platform for chat agents — workspaces, multi-provider LLMs, feedback corpus.
No known vulnerabilities
vulcn
v1.1.1 · 2 months ago
Security evals for the AI era. Probes · Targets · Graders · Proof. Confirmed XSS / SQLi / BOLA / prompt-injection / MCP-RCE with reproducible proof attached to every finding.
No known vulnerabilities
maestro-eval
v0.0.0 · 2 months ago
Reserved name — future eval harness (golden paths, injection, traffic replay) for the Maestro agent runtime. See https://github.com/costasoftware/maestro.
No known vulnerabilities
@trident-ai/cli
v0.4.0 · 4 days ago
Trident CLI — sign in, wire your repo into Trident with your coding agent, pentest what you deploy, and read findings from the terminal.
No known vulnerabilities
session-grep
v0.1.0 · 1 month ago
Grep AI coding-session transcripts (Claude Code, Codex) with bounded message context — built for agents. Includes rarity-ranked multi-word search, session overviews, sampled spines, and a built-in self-test.
No known vulnerabilities