evaluation - 技术专题

相关标签
open-sourceplaygroundmonitoringanalyticsevaluationself-hostedycombinatoropenaiobservabilityautogen

Here are 4,409 public repositories matching this topic...

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

  • Updated Jul 26, 2026
  • TypeScript
mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

  • Updated Jul 25, 2026
  • Python

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

  • Updated Jul 26, 2026
  • TypeScript