Model Merging Enables Cross-Domain Code Clone Detection
A recent study on arXiv investigates merging models to tackle fragmentation in code clone detection. Traditional detectors, often specialized, can experience F1 score declines exceeding 70% when used outside their intended conditions. The researchers adopted five task-vector techniques to merge parameters and conducted tests across four code models, three benchmarks, and twelve setups. Findings indicate that integrating similar types can yield successful cross-domain detectors, validated through two distinct model families and three random seeds. This research seeks to resolve complications arising from utilizing multiple specialized models and the difficulties of training a single detector with incomplete data.
Key facts
- Paper arXiv:2608.04215
- F1 drops exceeding 70% across domains
- Evaluates five task-vector methods
- Evaluates greedy layer stitching
- Evaluates cross-tokenizer alignment
- Four code models, three benchmarks, twelve configurations
- Same-base TIES merging effective
- Validated across two model families and three random seeds
Entities
Institutions
- arXiv