Skip to content

Does training on Claw-Anything data improve performance on Claw-Eval? #3

Description

@yaoshengyu

Hi, thanks for the great work!

The paper reports +23.7pp on the Claw-Anything eval set after fine-tuning.
We're wondering if training on this data also improves performance on Claw-Eval.

Given the large distribution gap (Claw-Anything has ~18x more context, 10x more services,
plus event logs that Claw-Eval doesn't have), we're concerned the trained model might
over-explore on simpler Claw-Eval tasks and hurt performance.

Did you run any cross-benchmark experiments? Any insight would be appreciated!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions