>100 Views
June 12, 26
スライド概要
佐藤 大允, 岡本 一志, 軽部 幸起, 原田 慧, 柴田 淳司: 大規模言語モデルを用いたトピックラベリング戦略と品質評価, 第40回人工知能学会全国大会, 2026.6, 群馬県高崎市.
Data Science Research Group, The University of Electro-Communications
2026.6.12 JSAI 2026 1 / 22
2026.6.12 JSAI 2026 2 / 22
2026.6.12 JSAI 2026 3 / 22
2026.6.12 JSAI 2026 4 / 22
LLM 2026.6.12 JSAI 2026 5 / 22
LLM [Rijcken+, 2023] [Alsulami+, 2025] LLM LLM LLM 2026.6.12 JSAI 2026 6 / 22
[Xu+, 2024] [Liu+, 2024] LLM Lost in the Middle 2026.6.12 JSAI 2026 7 / 22
LLM 2026.6.12 JSAI 2026 8 / 22
RQ LLM - 2026.6.12 JSAI 2026 9 / 22
2026.6.12 JSAI 2026 10 / 22
2022 Computer Science arXiv metadata Abstract arXiv Abstract Computer Science 22,901 171.6 words 170 words 12 words 464 words 2026.6.12 JSAI 2026 11 / 22
BERTopic Sentence Transformers LDA all-MiniLM-L6-v2 c-TF-IDF 10 LLM GPT-5 2026.6.12 JSAI 2026 12 / 22
2026.6.12 JSAI 2026 13 / 22
7 16 2026.6.12 JSAI 2026 14 / 22
- all-MiniLM-L6-v2, Qwen3-Embedding-0.6B, text-embedding-3-large 2026.6.12 JSAI 2026 15 / 22
Topic Label Condition 1 Graph Theory and Algorithms 1 Algorithms and computational complexity in graphs, networks, and automata 1 Algorithmic Graph Theory and Complexity 1 Combinatorics and computational complexity 1 Graph algorithms and computational complexity graphs / graph / vertex / polynomial / problem / vertices / algorithm / codes / number / log 2026.6.12 JSAI 2026 16 / 22
- 2026.6.12 JSAI 2026 17 / 22
- (a) all-MiniLM-L6-v2 2026.6.12 (b) Qwen3-Embedding-0.6B JSAI 2026 (c) text-embedding-3-large 18 / 22
Baseline Top Middle Bottom Balanced all-MiniLM-L6-v2 4.4 2.6 1.6 4.0 2.4 Qwen3-Embedding-0.6B 4.2 2.6 1.4 4.4 2.4 text-embedding-3-large 3.8 2.0 1.6 4.6 3.0 Overall Average Rank 4.1 2.4 1.5 4.3 2.6 1 2026.6.12 5 JSAI 2026 19 / 22
LLM RQ LLM 2026.6.12 JSAI 2026 20 / 22
arXiv LLM CS RQ LLM 2026.6.12 JSAI 2026 21 / 22
[Jung+, 2024] H. S. Jung, H. Lee, Y. S. Woo, S. Y. Baek, J. H. Kim: Expansive data, extensive model: Investigating discussion topics around LLM through unsupervised machine learning in academic papers and news, PLOS One, 19(5), 2024. [Liu+, 2025] J. Liu, Z. Shang, W. Ke, P. Wang, Z. Luo, J. Liu, G. Li, Y. Li: LLM-Guided Semantic-Aware Clustering for Topic Modeling, Proc. 63rd Annu. Meet. Assoc. Comput. Linguist., 2025. [Rijcken+, 2023] E. Rijcken, F. Scheepers, K. Zervanou, M. Spruit, P. Mosteiro, U. Kaymak: Towards Interpreting Topic Models with ChatGPT, Proc. 20th World Congr. Int. Fuzzy Syst. Assoc., 2023. [Alsulami+, 2025] M. M. Alsulami, M. A. Thafar: Enhancing Topic Interpretability with ChatGPT: A Dual Evaluation of Keyword and Context-Based Labeling, Int. J. Adv. Comput. Sci. Appl., 16(5), 2025. [Xu+, 2024] F. Xu, W. Shi, E. Choi: RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective Augmentation, Proc. 12th Int. Conf. Learn. Represent., 2024. [Grootendorst, 2022] M. Grootendorst: BERTopic: Neural topic modeling with a class-based TF-IDF procedure, arXiv preprint arXiv:2203.05794, 2022. [Liu+, 2024] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the Middle: How Language Models Use Long Contexts,” Trans. Assoc. Comput. Linguist., vol. 12, pp. 157–173, 2024. 2026.6.12 JSAI 2026 22 / 22