Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

By Ahmed El Kady · Paper · cs.LG

Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three d

Model Launch · Cs.lg

View original

HomeResourceLoading…