A comprehensive analysis of large language models in generative artificial intelligence-assisted research writing:insights from 2024 ASCO gastrointestinal oncology data by Chinese scholars
HAN Xu
LIU Liang
LOU Wen-hui
Abstract:Objective To delineate the current role of large language models(LLMs)in facilitating scientific research in oncology based on abstracts presented by Chinese scholars at ASCO in 2024.Methods 305 abstract papers on gastrointestinal cancer research published by Chinese scholars(including Hong Kong,Macao and Taiwan)at ASCO in 2024 were collected,and the generation probability of generative artificial intelligence(Gen AI)was detected by GPTZero Deep Learning.In 2021,60 papers of Chinese scholars without Gen AI tools were used as negative controls,and papers generated by Gen AI were used as positive controls.ROC curve evaluated the accuracy of GPTZero Deep Learning in detecting content generated by Gen AI.Pearson correlation analysis was used to verify the correlation between Gen AI generation probability and human creation probability,and the overall quality score was used to evaluate the quality of abstract papers.Results Among all abstracts included,Beijing(50 papers),Guangdong(49 papers)and Shanghai(48 papers)ranked the top three by region.The top five cancers were liver cancer(29.51%),esophageal cancer(18.69%),pan-cancer(14.43%),colorectal cancer(10.82%)and stomach cancer(10.82%).Immunotherapy remains the hottest topic,with PD-1,PD-L1,CTLA-4,PD-1/CTLA-4 and PD-1/TIGIT accounting for 77.28%,11.36%,2.27%,6.82%and 2.27%respectively.In molecular targeted drugs,tyrosine kinase inhibitors account for 85.19%,while multi-kinase inhibitors account for 11.11%.Under the optimal threshold of ROC curve,both the sensitivity and specificity of GPTZero Deep Learning to detect the accuracy of Gen AI content were 100%.GPTZero Deep Learning found that the probability of online abstract papers containing Gen AI content was higher in 2024 than that in 2021,with a statistically significant difference[17%(6%,35.5%)vs.5.5%(3%,12.75%),P<0.001].The probability of Gen AI generation under diagnostic confidence was negatively correlated with the probability of human creation(r=-0.852,P<0.001).The generation probability of Gen AI in eastern economic regions was higher than that in other provinces,and that in non-clinical study group was higher than that in clinical study group,with a statistically significant difference(P<0.05).The overall quality score of human-created abstracts was significantly higher than that of Gen AI generated,and the difference was statistically significant[(13.7±1.8)scores vs.(8.9±2.2)scores,P<0.001)].Conclusion Compared with 2021,the signal of AI content in ASCO abstracts of Chinese scholars in 2024 increased significantly,and the overall quality of human-created abstracts was better than that of Gen AI.
Keywords:artificial intelligencelarge language modelsAmerican Society of Clinical Oncologygastrointestinal cancerChinese scholars
Publication Date:2024-08-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:6( 894-899 )
Chinese Journal of Practical Surgery

Chinese Journal of Practical Surgery

ISTICPKUCSCD
ISSN:1005-2208
Year, Vol.(Issue):2024,44(8)