Skip to content
Kernelia
All news
ResearchHugging Face Blog

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

The post introduces ITBench-AA, the first benchmark aimed at measuring the performance of frontier AI models on enterprise IT tasks that involve agentic behavior. Results show that these models scored below 50 % on the benchmark, highlighting significant gaps in current capabilities.

Summary written by Kernelia from the original article by Hugging Face Blog. The story and its rights belong to its author.