Alibaba Cloud Bailian provides an application evaluation feature that supports both automated and manual evaluation methods to help you systematically assess and optimize the performance of agent applications.
Alibaba Cloud Bailian offers an application evaluation feature that enables systematic assessment of agent application output quality through two evaluation methods: automated evaluation and manual evaluation.
Bailian also provides a model evaluation feature. Unlike application evaluation, model evaluation is used to assess the foundational capabilities of individual models.
Evaluation Methods
Automated Evaluation
Leverages large language models to automatically generate evaluation datasets based on your knowledge base, then automatically evaluates agent responses and generates evaluation reports along with optimization recommendations. Supports both single-application evaluation and cross-application comparative evaluation.
Manual Evaluation
Involves manually constructing evaluation datasets and conducting human analysis and scoring of application responses to produce evaluation reports. Ideal for scenarios requiring domain expert judgment.
Related Features
- Evaluation Datasets — Store and manage data for evaluation tasks
- Evaluation Tasks — Create and execute evaluation tasks