We test the standard RLVR tool-use recipe (GRPO on Qwen2.5-7B-Instruct) on a deliberately minimal knowledge-graph tool API of four Freebase navigation verbs over Complex WebQuestions. The policy’s tool-grounded answer rate climbs from 3.8% to 9.6% before collapsing to 0% within a single 50-step window, a peak-then-collapse pattern replicated across four seeds and seven reward designs. We trace the failure to sparse interface feedback and show that one-iteration self-distillation reaches 40.0% exact-match accuracy, with the performance ceiling appearing interface-bound across the 7B-14B range.
@inproceedings{sun2026peakcollapse,title={Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use},author={Sun, Tianda and Kazakov, Dimitar},booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},address={Budapest, Hungary},year={2026},}
arXiv
Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams
Tool-using LLM agents produce trajectories whose calls form a directed dependency graph, where earlier tool outputs supply arguments to later calls. Using a low-capacity edge probe on the residual stream of Qwen3-32B, we decode this dependency graph well above both a Hewitt-Liang random-label control and a positional baseline, with counterfactual tests indicating the signal tracks abstract topology rather than identifier values. To our knowledge this is the first structural probe of an LLM agent’s runtime tool-call dependency graph, replicating across three multi-hop benchmarks and two model families.
@article{sun2026toolcall,title={Tool-Call Dependency Structure is Linearly Decodable in {LLM} Agent Residual Streams},author={Sun, Tianda and Kazakov, Dimitar},journal={arXiv preprint arXiv:2605.25310},year={2026},}
EMNLP
Kinship Data Benchmark for Multi-hop Reasoning
Tianda Sun and Dimitar Kazakov
In Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
We introduce KinshipQA, a contamination-proof benchmark evaluating LLM multi-hop reasoning across 7 culturally diverse kinship systems. We evaluate 6 state-of-the-art LLMs on 3,134 questions with controlled reasoning complexity (1-4 hops), revealing significant performance degradation as reasoning depth increases.
@inproceedings{sun2026kinshipqa,title={Kinship Data Benchmark for Multi-hop Reasoning},author={Sun, Tianda and Kazakov, Dimitar},booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},address={Budapest, Hungary},year={2026},}
2025
RANLP
KGEIR: Knowledge Graph-Enhanced Iterative Reasoning for Multi-Hop Question Answering
Tianda Sun and Dimitar Kazakov
In Proceedings of the RANLP 2025 Workshop on Robust and Low-Resource Large Language Models (R2LM), 2025
We propose KGEIR, a framework combining LLM reasoning with dynamic knowledge graph construction. KGEIR achieves competitive and superior performance on HotpotQA, 2WikiMultiHopQA, and MuSiQue benchmarks compared to state-of-the-art RAG methods, with a 4.5% F1 improvement over Tree-of-Thought baseline.
@inproceedings{sun2025kgeir,title={KGEIR: Knowledge Graph-Enhanced Iterative Reasoning for Multi-Hop Question Answering},author={Sun, Tianda and Kazakov, Dimitar},booktitle={Proceedings of the RANLP 2025 Workshop on Robust and Low-Resource Large Language Models (R2LM)},year={2025},address={Varna, Bulgaria},}
bioRxiv
TOGGLE Delineates Fate and Function within Individual Cell Types via Single Cell Transcriptomics
Jiawei Chen, Zhihao Chen, Tianda Sun, and 7 more authors
A cross-disciplinary collaboration combining computational algorithms with single cell transcriptomics to delineate cell fate and function within individual cell types.
@article{chen2025toggle,title={TOGGLE Delineates Fate and Function within Individual Cell Types via Single Cell Transcriptomics},author={Chen, Jiawei and Chen, Zhihao and Sun, Tianda and Liu, Kun and Jiang, Enming and Nong, Yongjie and Yuan, Tao and Dai, Chun-Chun and Yan, Yao and others},journal={bioRxiv},year={2025},doi={10.1101/2025.01.01.631041},}
2024
ECAI
A Hybrid Question Answering Model with Ontological Integration for Environmental Information
Tianda Sun, Jamie Carr, and Dimitar Kazakov
In Proceedings of DAO-XAI 2024 Workshop at ECAI, 2024
We present a hybrid question answering model that integrates ontological knowledge for environmental information retrieval, combining structured knowledge representation with neural language understanding.
@inproceedings{sun2024hybrid,title={A Hybrid Question Answering Model with Ontological Integration for Environmental Information},author={Sun, Tianda and Carr, Jamie and Kazakov, Dimitar},booktitle={Proceedings of DAO-XAI 2024 Workshop at ECAI},year={2024},address={Santiago de Compostela, Spain},}
This thesis investigates relation extraction from financial reports using distant supervision with BERT and PCNN architectures, leveraging the FIBO ontology for automatic training data generation.
@mastersthesis{sun2022relation,title={Relation Extraction from Financial Reports},author={Sun, Tianda},school={University of York},year={2022},}
IEEE
NLP Analysis of COVID-19 Radiology Reports in Indonesian using IndoBERT
Nunung Nurul Qomariyah, Tianda Sun, and Dimitar Kazakov
In IEEE International Biomedical Instrumentation and Technology Conference (IBIOMED), 2022
We apply IndoBERT for NLP analysis of COVID-19 radiology reports written in Indonesian, demonstrating the effectiveness of language-specific pre-trained models for medical text processing in low-resource languages.
@inproceedings{qomariyah2022nlp,title={NLP Analysis of COVID-19 Radiology Reports in Indonesian using IndoBERT},author={Qomariyah, Nunung Nurul and Sun, Tianda and Kazakov, Dimitar},booktitle={IEEE International Biomedical Instrumentation and Technology Conference (IBIOMED)},year={2022},note={British Council Newton Fund},}