ORCID iD
Garimella: https://orcid.org/0009-0004-7472-4690
Khandelwal: https://orcid.org/0000-0002-7291-3066
Document Type
Article
Abstract
Tool-calling evaluations often report whether the final call matches a reference answer, but this single outcome merges different failures: choosing a function outside the available registry, violating the selected schema, or filling arguments with values not supported by the user request. These failures matter more when models must choose from larger tool registries. We introduce RFCD, a registry-scaled function-calling benchmark built from BFCL, Glaive Function Calling v2, and Tool-Call-Data, with standardized JSON-schema registries ranging from 5 to 3,000 functions. We also introduce TAAG, a deterministic stage-decomposed evaluator for registry conformance, structural completeness, and argument grounding. Across six locally served sub-4B models, three data sources, and four feedback modes, registry scale is the dominant source of degradation: single-pass TAAG pass rate drops sharply from small to large registries on all sources. Repeated feedback improves some grounding metrics but adds substantial token cost and does not consistently improve end-to-end F1. Requirement-level TAAG-R feedback gives modest gains over stage-only TAAG-S feedback, which indicates that TAAG signals are useful for analysis and correction probes but do not solve registry scaling by themselves. RFCD and TAAG provide an evaluation resource for measuring where tool-calling systems fail as registries grow.
Publication Info
Spring 2026.
APA Citation
Garimella, R., Khandelwal, V., Kohli, A., & Sheth, A. (2026). RGFC: Registry Grounded Function Calling with stage-decomposed evaluation. Scholar Commons.
Rights
© 2026, The Authors