ResearchHugging Face Blog
ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM
The post introduces ITBench-AA, the first benchmark aimed at measuring the performance of frontier AI models on enterprise IT tasks that involve agentic behavior. Results show that these models scored below 50 % on the benchmark, highlighting significant gaps in current capabilities.
Summary written by Kernelia from the original article by Hugging Face Blog. The story and its rights belong to its author.
