verityResearch

Register

Results

Every entry below was read against pass/fail criteria that were written down and frozen before any result file was opened. The register lists studies in the order they resolved — favorable or not.

Published
01 study
Most recent
Protocol
Pre-registered · all entries
Code & license
github.com/verityResearch
Apache-2.0

Published results

Newest first · 1 entry

  1. /results/mcagent/

    Beyond Tool Access — mcAgent v20

    Verdict

    Fine-tuning earns its claim beyond tool access alone, on both evaluation sets.

    The base model essentially never emits a working tool call at all, regardless of phrasing. Tool access alone does not confer tool-use behavior at this model size.

    Pre-registered Fine-tuning Tool use Base vs. v20

    Held-out set — n = 106

    Tool-call rate
    0% rises to 98%
    Answer rate
    48% rises to 96%

    Adversarial probe — n = 43

    Tool-call rate
    2% rises to 83%
    Answer rate
    13% rises to 86%
    Read the study

    base model → fine-tuned v20

End of register

A study is added here only once its pre-registered criteria have been resolved. Work that has not reached that point is not listed.

Standard

What every entry on this register carries

Fixed before the run

  • Arm definitions. Which model, which tools, which prompts — specified in full for every arm.
  • Allowed interpretations. What counts as a valid reading of a response, written down so it cannot be widened afterwards.
  • Pass/fail criteria. The thresholds that decide the verdict, frozen alongside the arms.

Read after the run

  • The numbers, per arm and per evaluation set, with sample sizes stated.
  • A verdict that is whatever the frozen criteria say it is.
  • The code, under Apache-2.0, at github.com/verityResearch.
Questions
[email protected]