2026 · Research design, corpus construction & implementation
IDAgent: Bidirectional Idea–Dataset Discovery
We build a shared Idea–Dataset research memory for both idea-to-dataset retrieval and dataset-grounded research ideation.
- 217,942
- aligned Idea–Dataset pairs
- 64,252
- source papers
- 0.3334
- best relaxed MRR

Problem
Ideas describe goals and hypotheses while datasets describe fields, modalities, and empirical boundaries; keyword search cannot reliably align the two.
Approach
We constructed a large aligned corpus, trained an ID-SBERT/Faiss retriever, designed semantic clustering and cold-start evaluation, and implemented a RARO generation workflow.
Outcome
All 32 analysis tests pass, the 14-page paper recompiles, and 12 retrieval tables match generated CSVs. The interactive controller remains to be evaluated end to end.
