الملخص
Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Hossen, M., Selim, T., Gamgam, G., Yousif, T., Kasmi, A., Aissiou, I., Onipede, M., Butt, F. T., Zrigui, S., Paccotacya-Yanque, R. Y. G., Balayo, I., Elhouiti, I., Affes, H., Adhikari, B., Goyal, S., Isah, M. I., Bhat, M. I., Matia, S. K., Kadzue, P. K. M. T., Trabelsi, M., Owusu, E., Vinit, Majdoub, N., Alemnew, T., & Rekik, I. (2026). GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation. https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation
MLA 9
Hossen, Murad, et al. "GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation." https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation.
شيكاغو (المؤلف–التاريخ)
Hossen, Murad, Tasneem Selim, Gurur Gamgam, Tuga Yousif, Abderrahmane Kasmi, Ikram Aissiou, Mubaraq Onipede, Faran Taimoor Butt, Sanae Zrigui, Rosa Y. G. Paccotacya-Yanque, Ignatius Balayo, Ikram Elhouiti, Hadil Affes, Bijay Adhikari, Sargam Goyal, Muhammad Ibrahim Isah, Mohammad Idrees Bhat, Samuel Kangoni Matia, Peguy Kem-Meka Tiotsop Kadzue, Maha Trabelsi, Emmanuel Owusu, Vinit, Nour Majdoub, Tamiru Alemnew, and Islem Rekik. 2026. "GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation." https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation.
هارفارد
Hossen, M., Selim, T., Gamgam, G., Yousif, T., Kasmi, A., Aissiou, I., Onipede, M., Butt, F. T., Zrigui, S., Paccotacya-Yanque, R. Y. G., Balayo, I., Elhouiti, I., Affes, H., Adhikari, B., Goyal, S., Isah, M. I., Bhat, M. I., Matia, S. K., Kadzue, P. K. M. T., Trabelsi, M., Owusu, E., Vinit, Majdoub, N., Alemnew, T. and Rekik, I. (2026) 'GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation', Available at: https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation.
فانكوفر
Hossen M, Selim T, Gamgam G, Yousif T, Kasmi A, Aissiou I, et al. GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation. https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation
IEEE
M. Hossen, T. Selim, G. Gamgam, T. Yousif, A. Kasmi, I. Aissiou, M. Onipede, F. T. Butt, S. Zrigui, R. Y. G. Paccotacya-Yanque, I. Balayo, I. Elhouiti, H. Affes, B. Adhikari, S. Goyal, M. I. Isah, M. I. Bhat, S. K. Matia, P. K. M. T. Kadzue, M. Trabelsi, E. Owusu, Vinit, N. Majdoub, T. Alemnew, and I. Rekik, "GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation," https://omanscience.com/ar/articles/gnn-cb-a-graph-neural-network-competition-benchmark-for-human-and-llm-evaluation.