# llm-evaluation

このトピックのトレンドリポジトリ(3件)

langfuse/langfuse

langfuse/langfuseOtherTypeScript
25.9k

🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

analyticsautogenevaluationlangchainlarge-language-modelsllama-indexllmllm-evaluationllm-observabilityllmopsmonitoringobservabilityopen-sourceopenaiplaygroundprompt-engineeringprompt-managementself-hostedycombinator

AIアプリの品質とセキュリティを丸ごとテスト!GPT・Claude・Geminiを一括比較 — promptfoo

promptfoo/promptfooAITypeScript
15.6k3回登場

promptfooは、AIアプリ(ChatGPTのようなAIを使ったサービス)の品質チェックとセキュリティ検査を自動化するツールです。「この質問をしたらAIが正しく答えるか?」「悪意ある入力で情報が漏れないか?」といったテストを、設定ファイ

cici-cdcicdevaluationevaluation-frameworkllmllm-evalllm-evaluationllm-evaluation-frameworkllmopspentestingprompt-engineeringprompt-testingpromptsragred-teamingtestingvulnerability-scanners

Tencent/AI-Infra-Guard

Tencent/AI-Infra-GuardOtherPython
5.6k2回登場

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

agentagent-securityai-infraai-red-teamingai-securityllmllm-evaluationllm-jailbreakllm-securitymcp-scanopenclaw-securityprompt-injectionprompt-securityscannersecuritysecurity-toolsskill-scannerskills-securityvulnerability