← 返回
未分类 Key 中文

agent-evaluation

Evaluate LLM agents via behavioral tests, capability assessments, reliability metrics, and benchmarks to identify real-world performance issues.
通过行为测试、能力评估、可靠性指标和基准测试评估 LLM 代理,以识别真实世界的性能问题。
godferylindsay godferylindsay 来源
未分类 clawhub v1.0.2 1 版本 100000 Key: 需要
★ 0
Stars
📥 202
下载
💾 0
安装
1
版本
#ai#automation#latest

概述

Agent Evaluation

You're a quality engineer who has seen agents that aced benchmarks fail spectacularly in

production. You've learned that evaluating LLM agents is fundamentally different from

testing traditional software—the same input can produce different outputs, and "correct"

often has no single answer.

You've built evaluation frameworks that catch issues before production: behavioral regression

tests, capability assessments, and reliability metrics. You understand that the goal isn't

100% test pass rate—it

Capabilities

  • agent-testing
  • benchmark-design
  • capability-assessment
  • reliability-metrics
  • regression-testing

Requirements

  • testing-fundamentals
  • llm-fundamentals

API Integration

All LLM calls within agent evaluation workflows use SkillBoss API Hub via a single unified endpoint:

import requests, os

SKILLBOSS_API_KEY = os.environ["SKILLBOSS_API_KEY"]

def call_llm(messages: list) -> str:
    r = requests.post(
        "https://api.skillboss.co/v1/pilot",
        headers={"Authorization": f"Bearer {SKILLBOSS_API_KEY}", "Content-Type": "application/json"},
        json={
            "type": "chat",
            "inputs": {"messages": messages},
            "prefer": "balanced"
        },
        timeout=60,
    )
    return r.json()["result"]["choices"][0]["message"]["content"]

Required environment variables: SKILLBOSS_API_KEY

Patterns

Statistical Test Evaluation

Run tests multiple times and analyze result distributions

Behavioral Contract Testing

Define and test agent behavioral invariants

Adversarial Testing

Actively try to break agent behavior

Anti-Patterns

❌ Single-Run Testing

❌ Only Happy Path Tests

❌ Output String Matching

⚠️ Sharp Edges

IssueSeveritySolution
---------------------------
Agent scores well on benchmarks but fails in productionhigh// Bridge benchmark and production evaluation
Same test passes sometimes, fails other timeshigh// Handle flaky tests in LLM agent evaluation
Agent optimized for metric, not actual taskmedium// Multi-dimensional evaluation to prevent gaming
Test data accidentally used in training or promptscritical// Prevent data leakage in agent evaluation

Related Skills

Works well with: multi-agent-orchestration, agent-communication, autonomous-agents

版本历史

共 1 个版本

  • v1.0.2 当前
    2026-05-20 05:49 安全 安全

安全检测

腾讯云安全 (Keen)

安全,无风险
查看报告

腾讯云安全 (Sanbu)

安全,无风险
查看报告

🔗 相关推荐

design-media

video

godferylindsay
使用 SkillBoss API Hub 生成 MP4 视频,支持自定义时长、宽高比及可选图片输入,需提供 SKILLBOSS_API_KEY。
★ 0 📥 627
ai-agent

self-improving agent

pskoett
记录自身发现以实现自我改进的技能
★ 4,155 📥 928,200
ai-agent

Find Skills

root
帮助用户发现和安装智能体技能,当用户询问如「如何做X」、「找X的技能」、「有能做...的吗」等问题时
★ 1,512 📥 568,958