Document Type

Article

Abstract

Tool-calling evaluations often report whether the final call matches a reference answer, but this single outcome merges different failures: choosing a function outside the available registry, violating the selected schema, or filling arguments with values not supported by the user request. These failures matter more when models must choose from larger tool registries. We introduce RFCD, a registry-scaled function-calling benchmark built from BFCL, Glaive Function Calling v2, and Tool-Call-Data, with standardized JSON-schema registries ranging from 5 to 3,000 functions. We also introduce TAAG, a deterministic stage-decomposed evaluator for registry conformance, structural completeness, and argument grounding. Across six locally served sub-4B models, three data sources, and four feedback modes, registry scale is the dominant source of degradation: single-pass TAAG pass rate drops sharply from small to large registries on all sources. Repeated feedback improves some grounding metrics but adds substantial token cost and does not consistently improve end-to-end F1. Requirement-level TAAG-R feedback gives modest gains over stage-only TAAG-S feedback, which indicates that TAAG signals are useful for analysis and correction probes but do not solve registry scaling by themselves. RFCD and TAAG provide an evaluation resource for measuring where tool-calling systems fail as registries grow.

Rights

© 2026, The Authors

Share

COinS