レポート目次 Report contents 第2章 Chapter 2

「富岳サポートサイト」と「スパコン成果ナビ」生成AIチャットボットの質問クエリ分析と傾向比較 Query Analysis and Comparison of Trends for the Generative AI Chatbots on the Fugaku Support Site and the Supercomputer Report Navigator

2.1 クエリを分類することの意味

第1章では、質問数やセッション数の推移をもとに、「富岳サポートサイト」におけるAskDonaの利用状況を確認した。本章では、利用量から質問内容へと視点を移し、利用者がAskDonaに何を求めているのかを分析する。

RAGに求められる機能は、寄せられる質問の内容によって異なる。特定の仕様や数値を確認する質問であれば、該当する文書を正確に検索することが重要となる。一方、複数の研究事例を比較する質問や、条件に合う事例を一覧で求める質問では、複数の文書を横断し、情報を整理・集計する必要がある。また、利用者がエラーメッセージやプログラムを提示して原因を尋ねる場合は、登録された文書を検索するだけでは十分に答えられないこともある。

このように、AskDonaという同じRAG基盤を使用していても、チャットボットの用途が異なれば、寄せられる質問の型も異なる。質問の型が異なれば、回答に必要な検索方法や情報処理の仕組みも変わる。RAGの性能を一つの精度指標だけで捉えるのではなく、実際にどのような質問が寄せられ、それに対応する仕組みが備わっているかを確認する必要がある。

本章では、「富岳サポートサイト」と「スパコン成果ナビ」の2つのチャットボットを対象とする。「富岳サポートサイト」は、「富岳」を利用する過程で生じる手順の確認やトラブルへの対処を支援するものである。これに対して「スパコン成果ナビ」は、スーパーコンピュータを利用した研究成果を探し、研究テーマや事例を調査するために利用される。同じ分類体系で両者を比較することにより、チャットボットの用途が質問の構成とRAGに必要な機能へどのような違いを与えるのかを明らかにする。

「スパコン成果ナビ」とは - チャットボットの目的とデータベースの詳細

「スパコン成果ナビ」は、HPCI研究成果ページに登録・公開されている利用報告書などを、LLMとの対話を通じて横断的に調査するための研究成果閲覧支援サービスである。「富岳」だけでなく、「京」やHPCIを構成する国内の大学・研究機関の大規模計算機を用いて実施された研究、JHPCN(学際大規模情報基盤共同利用・共同研究拠点)の採択課題などを対象としている。

図2.1 スパコン成果ナビ イメージ
図2.1 スパコン成果ナビ イメージ

利用者は、知りたい研究テーマやキーワードを文章や単語で入力することで、関連する研究課題や成果を探すことができる。特定の課題を検索するだけでなく、複数の報告書の内容をまとめたり、共通点や違いを比較したりすることもできる。これまでにスーパーコンピュータがどのような研究に利用され、どのような成果が得られてきたのかを効率的に調べ、研究テーマの検討や関連事例の調査に役立てることを目的としている。

「スパコン成果ナビ」のデータベースには、HPCIに関する資料5,202件とJHPCNに関する資料1,356件の、合計6,558件のPDF文書が登録されている。対象年度は、HPCI関連資料が2012年度から2024年度、JHPCN関連資料が2011年度から2025年度である。各文書には、研究分野、利用枠、年度、課題番号などのメタデータ(文書を分類・検索するための付加情報)が設定されている。研究分野は、物質・材料・化学、工学・ものづくり、物理・素粒子・宇宙、情報・計算機科学・AI、バイオ・ライフ、環境・防災・減災など、10分野に分類されている。

登録されている資料には、研究の目的、手法、得られた成果をまとめた「成果概要」、その内容を簡潔に整理した「要約」、研究の過程や結論、発表論文などを含む「最終報告書」、研究内容を図やグラフとともに示した「研究紹介ポスター」などがある。そのため、研究分野や年度、利用した計算資源、課題番号などの明確な条件による検索に加え、「ある研究テーマに関連する事例を幅広く知りたい」「複数の研究成果を比較したい」といった探索的な質問にも利用できる。

このように、「スパコン成果ナビ」は、個別の操作方法やトラブルへの対処を支援する「富岳サポートサイト」とは異なり、蓄積された多数の研究成果から、条件に合う事例を探し、整理・比較することを主な目的としている。この用途の違いは、それぞれのチャットボットに寄せられる質問の型や、回答に必要となる検索・情報処理の方法にも表れる。

2.2 クエリ分類の考え方 - 先行研究の整理

検索システムや質問応答システムの研究では、利用者が入力したクエリをさまざまな観点から分類してきた。代表的な観点は、「利用者が何をしたいのか」「どのような答えを求めているのか」「システムが答えるためにどのような処理を必要とするのか」の3つである。

検索意図を分類した初期の研究として、Broder(2002)はウェブ検索をNavigational、Informational、Transactionalの3種類に整理した。Navigationalは特定のウェブサイトへ到達するための検索、Informationalは情報を得るための検索、Transactionalは購入やダウンロードなどの行為を行うための検索である。この分類は、検索が単一の行為ではなく、その背後に異なる目的があることを示した。

Rose and Levinson(2004)は、この考え方をさらに細分化した。特にInformationalを、特定の答えを求めるDirected、テーマについて広く知ろうとするUndirected、助言や手順を求めるAdvice、商品やサービスの所在を探すLocate、候補の一覧を求めるListに分けている。同じ「情報を得たい」という目的でも、単一の答えを求める質問と、候補を広く集める質問では、適切な検索結果の形が異なることを示している。スパコン成果ナビにおける「関連する研究を一覧で知りたい」という質問を考えるうえで、Listを独立した目的として扱うことは特に重要である。

情報探索の深さに注目した研究として、Marchionini(2006)は検索活動をLookup、Learn、Investigateに整理した。Lookupは既知の事実や特定の項目を確認する検索である。Learnは複数の情報を比較しながら理解を深める検索、Investigateは分析、統合、評価を伴う継続的な調査である。Marchioniniは、LearnとInvestigateを探索的検索に関係する活動として位置付けている。この区分は、研究テーマや関連事例を探しながら調査範囲を広げていく、研究者の情報探索を捉えるうえで有効である。

学術情報の検索には、一般的なウェブ検索とは異なる特徴もある。Li et al.(2017)は、学術検索サービスScienceDirectにおける3,900万件超のクエリログを分析した。その結果、検索結果が0件となるnull queryが全クエリの10.3%を占め、少なくとも1回のnull queryを含む検索セッションが全セッションの25.0%を占めることを示した。これは、研究者による検索では対象を名指しする検索と、テーマを広げながら探す検索の両方に対応する必要があることを示している。

質問が求める答えの形については、Bolotova et al.(2022)が、非factoid質問(単一の語句や数値だけでは答えられない質問)を6種類に分類している。6種類は、根拠に基づく説明を求めるEvidence-based、複数の対象を比べるComparison、経験に基づく助言を求めるExperience、理由を求めるReason、手順を求めるInstruction、複数の立場や論点を求めるDebateである。実際のチャットボットには、単純な一問一答だけでなく、理由、比較、手順、判断材料などを求める質問が寄せられる。したがって、質問のテーマだけでなく、利用者がどのような形の答えを期待しているかも分類する必要がある。

RAGに関する近年の研究では、クエリを「答えるために必要な処理」の観点から分類する試みが行われている。Jeong et al.(2024)のAdaptive-RAGは、質問の複雑性を、検索を必要としない質問、1回の検索で答えられる質問、複数回の検索と推論を必要とする質問の3段階に分けている。質問の難易度に応じて、検索を行わない方法、1回だけ検索する方法、検索と推論を繰り返す方法を使い分ける考え方である。

Zhao et al.(2024)は、外部データとの関係からクエリを4段階に整理している。文書に明示された事実を取り出す質問、複数の情報から暗黙の事実を導く質問、文書に記載された判断根拠を読み取る質問、複数の情報から明示されていない根拠を推論する質問である。これらの研究は、同じ情報探索に見える質問でも、単純な検索で答えられるものと、複数の情報を組み合わせなければ答えられないものがあることを示している。

一方、回答できない質問にも複数の原因がある。Barnett et al.(2024)は、RAGの失敗を7つの段階に整理した。(1) 参照文書に答えが存在しない、(2) 答えを含む文書が検索上位に入らない、(3) 取得した文書が回答生成時の入力に残らない、(4) 入力には答えがあるがAIが正しく抽出できない、(5) 指定された形式で回答できない、(6) 回答の詳しさが適切でない、(7) 必要な情報の一部が欠ける、という失敗である。この整理から、回答できなかったという結果だけでなく、どの段階で問題が起きたのかを区別する必要があることがわかる。

また、Larson et al.(2019)は、チャットボットが対応する目的の範囲に含まれない質問をout-of-scope(対応範囲外)として明示的に扱った。実際の運用では、利用者がチャットボットの対応範囲を完全に把握しているとは限らない。したがって、既存の分類に無理に当てはめるのではなく、対応範囲外を独立した分類として把握することが、誤った回答の防止とサービス範囲の見直しにつながる。

分類体系の作り方にも複数の方法がある。Tamkin et al.(2024)のClioは、あらかじめ決めたカテゴリを実利用ログへ当てはめるだけでなく、会話から話題や言語などの属性を抽出し、内容が近い会話をまとめることで、ボトムアップに分類階層を構築する。Chatterji et al.(2025)は、ChatGPTへの入力を、情報や助言を求めるAsking、成果物の作成や作業を求めるDoing、意見や感情を表すExpressingの3種類に分類した。同時に、実利用の約8割がPractical Guidance、Seeking Information、Writingという3つの話題に集中していることを示した。これらの研究は、実際の利用ログを分析し、利用実態に合わせて分類体系を検証・調整することの重要性を示している。

以上の先行研究から、クエリの分類には少なくとも2つの異なる視点が必要であることがわかる。第1は、利用者が何をしたいのかを捉える視点である。第2は、その質問へ答えるために、RAGがどのような検索や情報処理を必要とするのかを捉える視点である。本章では、前者を「目的軸」、後者を「機構軸」として分け、2つのチャットボットを共通の枠組みで分析する。

分類の観点 主な先行研究 本章への反映
検索・質問の目的 Broder (2002)、Rose and Levinson (2004)、Marchionini (2006)、Li et al. (2017) 利用者が何を知り、何を行いたいのかを示す「目的軸」
求める答えの形 Bolotova et al. (2022) 事実、手順、比較、理由、一覧などの区別
検索と推論の複雑性 Jeong et al. (2024)、Zhao et al. (2024) 回答に必要な検索・統合方法を示す「機構軸」と難易度
回答できない原因 Barnett et al. (2024)、Larson et al. (2019) 情報不足、検索失敗、対応範囲外などの区別
実利用ログからの分類 Tamkin et al. (2024)、Chatterji et al. (2025) 実際のクエリをもとに分類を調整・検証する方法

表2.2 先行研究と本章の分類の対応

2.3 本章で用いるデータセット・共通の分類体系・手法

本章で分析するデータの概要を表2.1に示す。2つのデータセットは、対象期間や質問数、チャットボットの利用目的が異なる。このため、単純な件数の大小だけではなく、各分類が全体に占める割合を中心に比較する。

項目 「富岳サポートサイト」 スパコン成果ナビ
主な用途 「富岳」の利用方法、仕様、エラーへの対処などの確認 研究テーマ、研究者、企業、課題、研究成果などの探索
対象期間 2024年7月1日〜2026年6月30日 2025年8月25日〜2026年6月30日
記録された質問数 12,285件 1,065件
分類対象となった質問数 主比較:層化標本1,065件(母集団12,285件を重み付き推定) 主比較:期間内全1,065件。感度分析:利用者が自発的に入力したクエリ(自発クエリ)827件
主な除外規則 主比較:標本1,065件全件(除外なし) 主比較は期間内全1,065件(運営側によるテスト質問および運用・動作確認用クエリ238件を含む)。感度分析ではテスト質問を除く自発クエリ827件を使用

表2.1 分析対象データの概要

本章では、利用者の目的と、回答に必要なRAGの処理を分けて捉えるため、4つの軸を用いてクエリを分類した。中心となるのは、目的軸Q(Questions)と機構軸M(Mechanism)である。

目的軸Qは、利用者が最終的に何を得たいのかを示す。Q1「手順・方法の把握」、Q2「トラブル対応」、Q3「事実・仕様の確認」、Q4「可否・条件・ポリシーの確認」、Q5「特定対象の参照」、Q6「探索的発見」、Q7「網羅・集計・一覧化」、Q8「概念理解・学習」、Q9「生成・作成の依頼」、Q10「判断・相談」、Q11「対話管理」の11種類から、各クエリに1つを付与した。機構軸Mは、そのクエリへ答えるために必要な検索・処理の型を示す。M1「1回の検索で1文書内の連続した箇所から答えられる(単一箇所参照)」、M2「1文書の複数箇所を統合する(単一文書内統合)」、M3「複数文書の横断や逐次検索を要する(複数文書・逐次検索)」、M5「利用者が提示したログやコードなどの診断・生成を要する(利用者持ち込み内容の診断・生成)」を区別した。さらに、M8「意図が曖昧で確認を要する(曖昧・要明確化)」、M9a「対象コーパスの範囲外である(原理的に対象外)」、M9b「守備範囲内だが文書が未整備である(守備範囲内・文書未整備)」、M10「会話操作や有人対応要求などの非情報(対話管理・非情報)」を設けた。複数の条件が重なる場合でも主コードは1つとし、M4「網羅・集約を要する」、M6「否定・除外条件を含む」、M7「時点や期間を扱う」は副フラグとして追加した。

これに加え、第1章との接続を確認するため、表層の文言・語句・キーワード分類に基づくK軸(Keyword)を用いた。K軸は、K1「方法探索」、K2「エラー・トラブル相談」、K3「情報探索」、K4「その他・分類外」の4種類である。また、回答が参照した文書数をC軸(Citation)とし、C0「引用なし」、C1「1〜2文書」、C2「3〜4文書」、C3「5文書以上」の4段階に分けた。K軸とC軸は機械的な規則で判定し、目的軸Qと機構軸Mはクエリ本文と会話文脈をもとにLLMによってクエリを分類する手法を採用した。なお、Q軸とM軸についてはLLMによる判定なので、再判定や追試において完全に合致するわけではない点は留意されたい。本手法の詳細な手順や実務的な工夫などは、「LLMを用いたクエリ分類を実務で行うには:AskDonaチャットボットの数万件のクエリ分析から得られた実践的知見」を参照してほしい。

最後に、機構軸Mから、RAGに必要な機能の段階を示すティアを算出した。T1は標準的なRAGで回答できるクエリ、T2は複数箇所・複数文書の検索や集約など検索側の拡張が必要なクエリ、T3は利用者が持ち込んだ内容の診断、対象範囲外、文書不足、意図の曖昧さなど、RAGの機構やデータベースの改善によって解決できうるクエリである。M1であっても、網羅・除外・時系列の副フラグが付く場合はT2とした。

「富岳サポートサイト」については、2024年7月1日から2026年6月30日までの母集団12,285件を、四半期とデータベース言語の組合せで16層に分け、1,065件を層化無作為抽出した。以下では、各層の母集団件数を標本件数で割った重みを用いて、母集団の構成比を推定する。スパコン成果ナビについては、2025年8月25日から2026年6月30日までの期間内全1,065件を主比較に用いる。このうち238件は運営側テスト質問であり、これを除く自発クエリ827件は感度分析および成果ナビ固有の傾向確認に用いる。対象期間と抽出方法が異なるため、2サイトの利用量や総件数ではなく、分類構成比を比較する。

「富岳サポートサイト」の1,065件はLLMが本判定を行い、その後、あらかじめ固定した人手QA標本198件を確認した。1,065件すべてを人手で再判定したものではない。レビュー結果は変更なし187件、補正11件、保留0件で、補正率は5.6%であった。補正11件は最終マスター、集計へ反映済みである。判定用アンカー20問の一致率は各バッチ90.0〜100.0%、平均92.3%であった。なお、アンカーの合格基準は当初の95%以上から90%以上へ事後変更しており、この点は方法上の留意点である。

2.4 「富岳サポートサイト」のクエリ構造

「富岳サポートサイト」では、利用者が実際の作業を進める過程で生じる質問が中心であった。 図2.4をみると、目的軸ではQ2「トラブル対応」が22.7%で最も多く、Q1「手順・方法の把握」が21.3%で続いた。この2分類だけで44.1%を占める。さらに、Q3「事実・仕様の確認」が12.8%、Q4「可否・条件・ポリシーの確認」が9.9%、Q8「概念理解・学習」が9.5%であった。これに対し、Q6「探索的発見」は1.6%、Q7「網羅・集計・一覧化」は0.8%にとどまった。「富岳サポートサイト」では、まだ知らない事例を広く探すというより、計算機を利用するための具体的な手順、エラーの解消、仕様や条件の確認が求められている。

図2.4 目的軸(Q軸)の構成比の比較
図2.4 目的軸(Q軸)の構成比の比較

なお、LLMによる判定の妥当性の検証および1章とのつながりを担保するために、K軸による分類も同時に行っている。その結果、K軸では「その他・分類外」とされていたクエリの多くが、Q軸を用いることで利用者の具体的な目的(Q1〜Q11)へ適切に分類できることが確認できた。

機構軸では、M5「利用者持ち込み内容の診断・生成」が35.1%で最も多かった。利用者がエラーメッセージ、ジョブスクリプト、コマンド、数値などを提示し、その内容に即した原因の推定や修正案を求める質問が多いことを示している。M1「単一箇所参照」は27.6%、M2「単一文書内統合」は17.7%であり、既存文書から回答できる質問も一定数を占めた。一方、M3「複数文書・逐次検索」は8.0%であった。

富岳側の主要な推定値には標本誤差があるが、M5の95%信頼区間は32.4〜37.8%、M1は25.0〜30.2%であり、両者が主要な型であるという傾向は変わらない。目的軸でも、Q2の95%信頼区間は20.3〜25.1%、Q1は19.0〜23.7%であり、手順とトラブル対応が中心であることが確認できる。

難易度ティアでは、T3「機構やデータベースの改善で解決」が46.8%で最も多く、T2「検索側の拡張が必要」が27.6%、T1「標準的なRAGで解ける」が25.6%であった。T3が多い最大の理由はM5「利用者持ち込み内容の診断・生成」が35.1%で最大を占めるからである。

以上の結果をまとめると、以下のことがいえる。すなわち、「富岳サポートサイト」の改善は検索精度だけに限定するのではなく、文書検索に加え、利用者が提示したログやコードを安全に解析する仕組み、追加情報を尋ねる対話、必要に応じて有人対応へつなぐ導線が重要となる。こうした改善策は、日夜AskDonaが「富岳サポートサイト」を運営する中で、取り組んでいることである。

2.5 「スパコン成果ナビ」のクエリ構造

スパコン成果ナビでは、運営側テスト質問を除く自発クエリ827件を用いて、利用者自身が入力した質問の特徴を確認した。目的軸では、Q6「探索的発見」が32.6%で最も多く、Q7「網羅・集計・一覧化」が23.7%で続いた。両者を合わせると56.3%となる。Q5「特定対象の参照」も18.6%を占めた。利用者は、特定の研究課題を名指しして探すだけでなく、分野やテーマから関連事例を見つけたり、条件に合う複数の課題を一覧や表として取得したりする目的で「スパコン成果ナビ」を利用している。

図2.5-1をみると、機構軸ではM3「複数文書・逐次検索」が77.0%を占めた。M4「網羅・集約」の副フラグも29.7%に付与され、M7「時系列・鮮度」は6.7%であった。これは、単に関連文書を1件検索するだけでなく、複数の研究課題や成果報告書を横断し、条件に合う対象を絞り込み、一覧化・比較・集計する処理が頻繁に必要であることを示す。

図2.5-1 機構軸(M主コード)の構成比の比較
図2.5-1 機構軸(M主コード)の構成比の比較

図2.5-2で示した難易度ティアでは、T2「検索側の拡張が必要」が79.2%を占め、T1「標準的なRAGで解ける」は0.7%にすぎなかった。T3「機構やデータベースの改善で解決」は20.1%であり、そのうちM9a「原理的に対象外」は11.1%、M9b「守備範囲内・文書未整備」は1.2%であった。M9aとM9bは、いずれもそのままでは回答しにくいが、改善方法は異なる。M9aには対応範囲を明示し、当該チャットボットの回答範囲外であること、適切な情報源へ誘導することなどが必要である。M9bには、利用者が求める情報を特定し、文書やメタデータを追加することが有効である。

図2.5-2 難易度ティアの構成比の比較
図2.5-2 難易度ティアの構成比の比較

引用分類でも、複数文書を前提とする特徴が表れた。回答が存在する自発クエリ804件のうち、5文書以上を参照したC3は74.6%であった。運営側テスト質問を含む全1,065件と自発クエリ827件を比べると、Q7「網羅・集計・一覧化」は31.5%から23.7%へ低下し、Q6「探索的発見」は28.4%から32.6%へ上昇した。テスト質問が一覧化の割合を押し上げているものの、これを除いても、探索的発見と網羅・一覧化が主要な目的であり、M3「複数文書・逐次検索」が大半を占めるという全体像は変わらない。

なお、両サイトでM軸とQ軸それぞれの分類に該当する質問の具体例は、「LLMを用いたクエリ分類を実務で行うには:AskDonaチャットボットの数万件のクエリ分析から得られた実践的知見」に記載している。

2.6 両サイトの比較 - 用途がクエリの型を決める

双方1,065件で比較すると、目的軸の構成は大きく異なった。「富岳サポートサイト」ではQ1「手順・方法の把握」が21.3%、Q2「トラブル対応」が22.7%であるのに対し、「スパコン成果ナビ」ではそれぞれ8.3%、0.5%であった。反対に、「スパコン成果ナビ」ではQ6「探索的発見」が28.4%、Q7「網羅・集計・一覧化」が31.5%であるのに対し、「富岳サポートサイト」ではそれぞれ1.6%、0.8%であった。「富岳サポートサイト」では既知のシステムを利用するための支援が求められ、「スパコン成果ナビ」では未知の研究事例を発見し、複数の候補を整理する支援が求められている。この用途の違いが、クエリの目的構成と関連していると考えられる。

機構軸の差はさらに大きい。「スパコン成果ナビ」ではM3「複数文書・逐次検索」が75.9%であるのに対し、「富岳サポートサイト」では8.0%であり、67.9ポイントの差があった。一方、「富岳サポートサイト」ではM5「利用者持ち込み内容の診断・生成」が35.1%、M1「単一箇所参照」が27.6%であるのに対し、「スパコン成果ナビ」ではそれぞれ1.2%、0.6%であった。同じRAG基盤を用いていても、「スパコン成果ナビ」では複数文書検索と集約が中心となり、「富岳サポートサイト」では正確な単一箇所検索と、文書外の入力を扱う診断・生成の双方が必要となる。

クエリの難易度のティア比較でも、「スパコン成果ナビ」はT2が83.4%を占めるのに対し、「富岳サポートサイト」はT1が25.6%、T2が27.6%、T3が46.8%に分かれた。「スパコン成果ナビ」の中心課題は、検索対象を複数文書へ広げ、条件検索、再検索、集計、順位付け、時系列処理を組み合わせることである。「富岳サポートサイト」の中心課題は、単純な文書参照で回答できる質問を確実に処理しつつ、ログやコードを含む質問を別の処理系へ振り分けることである。したがって、両チャットボットへ同じ検索戦略と評価指標を一律に適用するのではなく、クエリの型に応じて処理経路を切り替える設計が必要である。

引用機能が安定した2025年10月から2026年6月の共通期間で、引用文書数別に回答数を集計したものが図2.6である。5文書以上を参照した割合は、「富岳サポートサイト」が38.0%、「スパコン成果ナビ」が72.3%であった。「富岳サポートサイト」では1〜2文書が27.1%、3〜4文書が32.4%で、両者を合わせて59.5%となる。なお、C軸では回答欠損を分母から除外している。「スパコン成果ナビ」では5文書以上の参照が中心であるのに対し、「富岳サポートサイト」では少数の文書を組み合わせる回答も多い。この結果は、「スパコン成果ナビ」でM3「複数文書・逐次検索」が多く、「富岳サポートサイト」でM1「単一箇所参照」とM2「単一文書内統合」が相対的に多いことと整合する。この背景には、「スパコン成果ナビ」のデータベースに登録されているドキュメントが比較的短く、件数が多いのに対して、「富岳サポートサイト」では1ドキュメント数百ページにわたるような、実機の操作マニュアルなどが限られた数で登録されているという、両データベースの相違点も要因として指摘できる。

図2.6 引用文書数(C軸)の構成比の比較(2025年10月〜2026年6月)
図2.6 引用文書数(C軸)の構成比の比較(2025年10月〜2026年6月)

なお、以上の比較は、用途と分類構成の関連を示すものであり、用途が差を直接引き起こしたと断定するものではない。対象期間、母集団、運営側テスト質問の有無、標本抽出の方法が異なる点には留意が必要である。ただし、「スパコン成果ナビ」のテスト質問を除く自発クエリでも主要な傾向が維持され、「富岳サポートサイト」の主要推定値の差も信頼区間より十分に大きい。少なくとも、両チャットボットに完全に同じRAG機能だけを実装すれば、それぞれの利用者のニーズを満たすのに十分であるとは考えにくい。

2.7 本章のまとめ

本章では、「富岳サポートサイト」と「スパコン成果ナビ」のクエリを、利用者の目的と回答に必要な機構に分けて比較した。その結果、次の3点が明らかになった。

第一に、「富岳サポートサイト」では、手順の確認とトラブル対応が中心であり、利用者が提示したログやコードを診断する質問が多かった。既存文書の検索精度を高めるだけでなく、入力内容の解析、追加質問、有人対応を含む支援が必要である。

第二に、「スパコン成果ナビ」では、探索的発見と網羅・一覧化が中心であり、複数文書・複数課題を横断する質問が大半を占めた。検索結果を並べるだけでなく、条件による絞り込み、集約、比較、時系列処理を行う機能が重要である。

第三に、同じRAG基盤を用いるチャットボットでも、用途の違いに応じて、クエリの型と必要な機能の構成は大きく異なっていた。RAGの評価では、全体の正答率だけでなく、クエリを型別に分け、各型に必要な検索・統合・診断・誘導が実行できたかを確認する必要がある。

今後は、M9a「原理的に対象外」、M9b「守備範囲内・文書未整備」などの境界事例を継続的に監査する必要がある。さらに、各分類と回答品質、利用者評価、再質問の有無を結び付けることで、どのクエリ型で失敗が生じやすいかを検証することが課題となる。

参考文献

総括

本レポートの位置づけ

本レポートでは、スーパーコンピュータ「富岳」のサポートサイトにおける生成AIチャット「AskDona」について、導入から24か月間の運用実績と、実際に寄せられたクエリの構造分析を報告した。2025年版レポートが「導入と、その効果の検証」を扱ったものであるとすれば、本レポートは「利用の定着と、クエリの中身の解明」を扱ったものと位置づけられる。

24か月にわたる利用データから、生成AIサービスが導入直後の一時的な利用にとどまらず、実務の中で継続して利用されている状況を確認できた。また、この期間中に実施した回答生成エンジンのアップデートと、その前後における利用状況の変化を追跡できたことも、実運用データならではの成果である。こうした長期的な定量データは、エンタープライズ領域における生成AI活用を検討する企業や研究機関にとって、実践的な参考事例になると考えられる。

本レポートで明らかになったこと

第1章では、24か月間に13,461件の質問が寄せられ、月ごとの変動を伴いながらも利用が継続・拡大していることを確認した。回答文字数の中央値は約2.8倍に増加しており、回答内容の充実によって、より少ないやり取りで疑問を解消できるケースが増えた可能性が示唆された。また、有人問い合わせにあたる「質問」チケットは、AskDona導入後2年目に1年目比43.6%、導入前比58.9%減少しており、AskDonaが有人窓口を補完する受け皿として機能している可能性が示された。

第2章では、「富岳サポートサイト」と「スパコン成果ナビ」という、用途の異なる2つのチャットボットに寄せられたクエリを、共通の分類体系で比較した。その結果、同じRAG基盤を用いていても、用途によって必要とされる機能の構成が大きく異なることが明らかになった。

「富岳サポートサイト」では、利用者が提示したログやコードの診断を必要とする質問が最も多かった。一方、「スパコン成果ナビ」では、複数の文書を横断して検索し、その結果を集約する処理が大半を占めた。これらの結果から、RAGの性能を単一の正答率だけで評価することには限界があり、クエリの型に応じて処理経路を切り替える設計が必要であることを、実データに基づいて確認できた。

今後の展望

今後は、本レポートで明らかになったクエリ構造をもとに、それぞれの型に適した処理を選択できるRAGアーキテクチャの高度化を進める。

とりわけ、利用者が提示したログやコードなどの内容を診断する機能、対応範囲外の質問を適切な情報源や有人窓口へ誘導する機能、必要な文書が不足している領域を継続的に特定・補充する仕組みが重要となる。これらはいずれも、技術開発と運用改善の両面から取り組むべき課題である。

当社は、R-CCSとの協力関係を継続・発展させ、こうした課題に引き続き取り組んでいく。本レポートで公開した知見が、同様の課題を抱える企業や研究機関にとって有益な示唆となり、日本の科学技術とAI活用の発展、さらには生成AI技術の社会実装の促進に貢献できることを期待している。

謝辞

本プロジェクトの遂行および本報告書の作成にあたり、理化学研究所計算科学研究センターの皆様から格別のご支援とご助言を賜りました。ここに記し、深く感謝申し上げます。

なお、本報告書に示す見解および結論は、すべて当社のものであり、上記所属機関の公式見解を示すものではありません。

2.1 Why Classifying Queries Matters

Chapter 1 examined the use of AskDona on the Fugaku Support Site based on trends in questions and sessions. This chapter shifts the focus from usage volume to question content and analyzes what users seek from AskDona.

The capabilities required of a RAG system depend on the content of the questions it receives. For a question asking about a specific specification or numerical value, accurately retrieving the relevant document is essential. By contrast, a question comparing multiple research cases or asking for a list of cases that meet specified conditions requires the system to search across documents and organize or aggregate information. When a user provides an error message or program and asks about its cause, searching the registered documents alone may not be sufficient to answer the question.

Thus, even when two chatbots use the same AskDona RAG platform, differences in their purposes lead to differences in the types of questions they receive. Different question types, in turn, require different retrieval methods and information-processing mechanisms. RAG performance should not be evaluated using only a single accuracy metric. It is also necessary to examine the types of questions received and whether the system includes mechanisms capable of handling them.

This chapter examines two chatbots: the Fugaku Support Site and the Supercomputer Report Navigator. The Fugaku Support Site assists users with procedures and problems arising in the course of using Fugaku. The Supercomputer Report Navigator is used to find research outcomes produced using supercomputers and to investigate research themes and related cases. By comparing the two under the same classification framework, we identify how chatbot purpose is associated with differences in query structure and in the capabilities required of RAG.

What Is the Supercomputer Report Navigator? Purpose of the Chatbot and Details of Its Database

The Supercomputer Report Navigator is a research-outcome browsing support service that enables users to search across usage reports and other materials registered and published on the HPCI Research Outcomes website through dialogue with an LLM. It covers not only Fugaku but also research conducted using the K computer and large-scale computing systems at Japanese universities and research institutions participating in HPCI, as well as projects selected by JHPCN (Joint Usage/Research Center for Interdisciplinary Large-scale Information Infrastructures).

Figure 2.1. Supercomputer Report Navigator
Figure 2.1. Supercomputer Report Navigator

Users can find relevant research projects and outcomes by entering a research topic or keywords as a sentence or individual terms. In addition to searching for a specific project, they can summarize content from multiple reports and compare commonalities and differences. The service is intended to help users efficiently investigate how supercomputers have been used in research and what outcomes have been obtained, supporting the development of research themes and searches for related cases.

The Supercomputer Report Navigator database contains 6,558 PDF documents: 5,202 HPCI-related materials and 1,356 JHPCN-related materials. The HPCI materials cover fiscal years 2012 through 2024, and the JHPCN materials cover fiscal years 2011 through 2025. Each document has metadata, or supplementary information used to classify and search documents, including research field, allocation category, fiscal year, and project number. Research is classified into ten fields, including materials and chemistry; engineering and manufacturing; physics, particle physics, and astronomy; information and computer science and AI; biological and life sciences; and environment, disaster prevention, and disaster mitigation.

The registered materials include Research Outcome Summaries describing the purpose, methods, and results of a study; Abstracts providing concise summaries; Final Reports covering the research process, conclusions, and related publications; and Research Introduction Posters presenting research with figures and graphs. The service can therefore be used not only for searches based on explicit conditions such as research field, fiscal year, computing resource used, and project number, but also for exploratory questions such as "I want to learn broadly about cases related to a certain research theme" and "I want to compare multiple research outcomes."

In this way, unlike the Fugaku Support Site, which assists users with specific operating procedures and troubleshooting, the Supercomputer Report Navigator is primarily intended to find, organize, and compare cases that meet specified conditions across a large body of accumulated research outcomes. This difference in purpose is also reflected in the types of questions received by each chatbot and in the retrieval and information-processing methods required to answer them.

2.2 Approach to Query Classification: Review of Prior Research

Research on search systems and question-answering systems has classified user queries from a variety of perspectives. Three representative perspectives are what the user wants to do, what form of answer the user seeks, and what processing the system requires in order to answer.

In an early study of search intent, Broder (2002) classified web searches as Navigational, Informational, or Transactional. Navigational searches seek a particular website, Informational searches seek information, and Transactional searches seek to perform an action such as making a purchase or downloading a file. This classification demonstrated that search is not a single activity, but serves different underlying purposes.

Rose and Levinson (2004) further subdivided this concept. In particular, they divided Informational searches into Directed, seeking a specific answer; Undirected, seeking broad knowledge about a topic; Advice, seeking advice or instructions; Locate, seeking the location of a product or service; and List, seeking a set of candidates. Even when the common purpose is to obtain information, a question seeking a single answer requires a different form of search result from one seeking a broad set of candidates. Treating List as an independent purpose is particularly important when considering questions such as "Show me a list of related research" in the Supercomputer Report Navigator.

Focusing on the depth of information seeking, Marchionini (2006) organized search activities into Lookup, Learn, and Investigate. Lookup confirms a known fact or specific item. Learn deepens understanding by comparing multiple pieces of information, while Investigate is a continuing inquiry involving analysis, synthesis, and evaluation. Marchionini positioned Learn and Investigate as activities related to exploratory search. This distinction is useful for describing how researchers expand the scope of an investigation while searching for research themes and related cases.

Academic information retrieval also has characteristics that differ from general web search. Li et al. (2017) analyzed more than 39 million query logs from the academic search service ScienceDirect. They found that null queries, which returned no search results, accounted for 10.3% of all queries, and that sessions containing at least one null query accounted for 25.0% of all sessions. These results indicate that systems supporting researchers must handle both searches for explicitly named targets and searches that broaden a topic during exploration.

Regarding the form of answer sought, Bolotova et al. (2022) classified non-factoid questions, which cannot be answered with a single term or number, into six types: Evidence-based, seeking an evidence-based explanation; Comparison, comparing multiple targets; Experience, seeking advice based on experience; Reason, seeking a reason; Instruction, seeking a procedure; and Debate, seeking multiple viewpoints or issues. In practice, chatbots receive questions seeking not only simple question-and-answer responses but also reasons, comparisons, procedures, and information needed for judgment. Queries therefore need to be classified not only by topic but also by the form of answer the user expects.

Recent RAG research has also attempted to classify queries according to the processing required to answer them. Adaptive-RAG, proposed by Jeong et al. (2024), divides question complexity into three levels: questions requiring no retrieval, questions answerable with one retrieval, and questions requiring multiple rounds of retrieval and reasoning. The approach selects among no retrieval, a single retrieval, or iterative retrieval and reasoning according to question complexity.

Zhao et al. (2024) organized queries into four levels based on their relationship to external data: questions that extract facts explicitly stated in documents; questions that derive implicit facts from multiple pieces of information; questions that identify the rationale for a judgment stated in documents; and questions that infer an unstated rationale from multiple pieces of information. These studies show that even questions that appear to involve the same type of information seeking may include both questions answerable through simple retrieval and questions that require combining multiple pieces of information.

Questions that cannot be answered also fail for multiple reasons. Barnett et al. (2024) organized RAG failures into seven points: (1) the answer is absent from the reference documents; (2) a document containing the answer is not ranked among the top retrieval results; (3) the retrieved document does not remain in the input used for answer generation; (4) the answer is present in the input but the AI fails to extract it correctly; (5) the system fails to answer in the specified format; (6) the level of detail is inappropriate; and (7) part of the necessary information is missing. This framework shows the need to distinguish the stage at which a problem occurs rather than considering only the outcome that a question could not be answered.

Larson et al. (2019) explicitly treated questions outside the range of intents handled by a chatbot as out of scope. In actual operation, users do not necessarily have a complete understanding of a chatbot's scope. Instead of forcing such questions into an existing category, identifying out-of-scope questions as a separate class can help prevent incorrect answers and support reviews of the service's scope.

There are also multiple approaches to constructing a classification framework. Clio, developed by Tamkin et al. (2024), does more than apply predetermined categories to real-world usage logs. It extracts attributes such as topic and language from conversations and groups similar conversations to build a classification hierarchy from the bottom up. Chatterji et al. (2025) classified ChatGPT inputs into three types: Asking, seeking information or advice; Doing, requesting the production of an artifact or the performance of a task; and Expressing, conveying opinions or emotions. They also found that approximately 80% of real-world use was concentrated in three topics: Practical Guidance, Seeking Information, and Writing. These studies demonstrate the importance of analyzing actual usage logs and validating and adjusting a classification framework to reflect real-world use.

Taken together, these studies show that query classification requires at least two distinct perspectives. The first captures what the user wants to do. The second captures the retrieval and information processing that a RAG system requires to answer the question. In this chapter, we separate these as the Purpose axis and the Mechanism axis, respectively, and use a common framework to analyze the two chatbots.

Classification perspective Key prior research Application in this chapter
Purpose of the search or question Broder (2002), Rose and Levinson (2004), Marchionini (2006), Li et al. (2017) Purpose axis representing what the user wants to know or do
Form of answer sought Bolotova et al. (2022) Distinction among facts, procedures, comparisons, reasons, lists, and other answer forms
Complexity of retrieval and reasoning Jeong et al. (2024), Zhao et al. (2024) Mechanism axis and difficulty level representing the retrieval and synthesis required to answer
Reasons a question cannot be answered Barnett et al. (2024), Larson et al. (2019) Distinction among insufficient information, retrieval failure, out-of-scope questions, and other causes
Classification based on real-world usage logs Tamkin et al. (2024), Chatterji et al. (2025) Methods for adjusting and validating classifications based on actual queries

Table 2.2. Relationship Between Prior Research and the Classification Used in This Chapter

2.3 Datasets, Common Classification Framework, and Methods Used in This Chapter

Table 2.1 summarizes the data analyzed in this chapter. The two datasets differ in study period, number of questions, and intended use of the chatbot. We therefore compare primarily the share of each classification in the relevant dataset, rather than simply comparing absolute counts.

Item Fugaku Support Site Supercomputer Report Navigator
Primary use Confirming how to use Fugaku, its specifications, and how to address errors Exploring research themes, researchers, companies, projects, research outcomes, and related information
Study period July 1, 2024 to June 30, 2026 August 25, 2025 to June 30, 2026
Recorded questions 12,285 1,065
Questions included in classification Primary comparison: stratified sample of 1,065 questions (weighted estimates for a population of 12,285) Primary comparison: all 1,065 questions in the period. Sensitivity analysis: 827 queries entered voluntarily by users (spontaneous queries)
Primary exclusion rules Primary comparison: all 1,065 questions in the sample (no exclusions) Primary comparison: all 1,065 questions in the period, including 238 test questions submitted by the operator and queries used for operational and functionality checks. Sensitivity analysis: 827 spontaneous queries after excluding test questions

Table 2.1. Overview of the Data Analyzed

To distinguish the user's purpose from the RAG processing required for an answer, we classified queries along four axes. The central axes are Purpose Q (Questions) and Mechanism M.

Purpose Q represents what the user ultimately wants to obtain. Each query received one of 11 primary codes: Q1 "Understanding procedures and methods," Q2 "Troubleshooting," Q3 "Confirming facts and specifications," Q4 "Confirming feasibility, conditions, and policies," Q5 "Locating a specific target," Q6 "Exploratory discovery," Q7 "Comprehensive retrieval, aggregation, and listing," Q8 "Conceptual understanding and learning," Q9 "Requesting generation or creation," Q10 "Decision support and consultation," and Q11 "Conversation management." Mechanism M represents the type of retrieval and processing required to answer the query. We distinguished M1 "Answerable from a contiguous passage in one document after one retrieval (single-passage lookup)," M2 "Integrates multiple passages in one document (within-document synthesis)," M3 "Requires cross-document or iterative retrieval (cross-document and iterative retrieval)," and M5 "Requires diagnosis or generation based on logs, code, or other content supplied by the user (diagnosis or generation from user-provided content)." We also defined M8 "The intent is ambiguous and requires clarification (ambiguous or clarification required)," M9a "Outside the target corpus in principle (intrinsically out of scope)," M9b "Within the service scope but the documentation is incomplete (in scope, documentation incomplete)," and M10 "Non-informational content such as conversation control or a request for human assistance (conversation management or non-informational)." Even when multiple conditions applied, each query received one primary code. M4 "Requires comprehensive retrieval or aggregation," M6 "Contains negation or exclusion conditions," and M7 "Involves a point in time or time period" were added as secondary flags.

In addition, to connect the analysis with Chapter 1, we used a Keyword axis K based on surface wording, phrases, and keywords. The K axis has four categories: K1 "Method seeking," K2 "Error and troubleshooting consultation," K3 "Information seeking," and K4 "Other or unclassified." We also defined a Citation axis C based on the number of documents cited in an answer: C0 "No citations," C1 "1-2 documents," C2 "3-4 documents," and C3 "5 or more documents." The K and C axes were assigned using mechanical rules. For the Purpose Q and Mechanism M axes, we adopted an LLM-based method that classified queries using the query text and conversational context. Because Q- and M-axis judgments are made by an LLM, results may not match perfectly when queries are reclassified or the procedure is repeated. For detailed procedures and practical techniques, see LLM-based Query Classification Practices: Lessons from Analyzing Tens of Thousands of AskDona Chatbot Queries.

Finally, we derived tiers from Mechanism M to indicate the level of RAG capability required. T1 consists of queries answerable by standard RAG. T2 consists of queries requiring retrieval-side extensions, such as retrieval from multiple passages or documents and aggregation. T3 consists of queries involving diagnosis of user-provided content, out-of-scope topics, insufficient documentation, or ambiguous intent that may be addressed through improvements to the RAG mechanism or database. Even an M1 query was assigned to T2 when it carried a secondary flag for comprehensive retrieval, exclusion, or temporal conditions.

For the Fugaku Support Site, the population of 12,285 questions from July 1, 2024, through June 30, 2026, was divided into 16 strata defined by the combination of quarter and database language, and 1,065 questions were selected through stratified random sampling. The population shares reported below are estimated using a weight calculated by dividing the population count in each stratum by its sample count. For the Supercomputer Report Navigator, the primary comparison uses all 1,065 questions recorded from August 25, 2025, through June 30, 2026. Of these, 238 were operator-submitted test questions. The remaining 827 spontaneous queries are used for sensitivity analysis and to examine patterns specific to the Supercomputer Report Navigator. Because the study periods and sampling methods differ, we compare the shares of classifications rather than usage volumes or total counts between the two sites.

An LLM assigned the primary classifications for the 1,065 Fugaku Support Site questions, after which a predetermined human-QA sample of 198 questions was reviewed. The review did not reclassify all 1,065 questions manually. Of the 198 reviewed questions, 187 were unchanged, 11 were corrected, and none were left pending, for a correction rate of 5.6%. The 11 corrections have been incorporated into the final master data and aggregates. Agreement on the 20 anchor questions used for classification ranged from 90.0% to 100.0% across batches, with an average of 92.3%. The anchor acceptance threshold was changed after the fact from at least 95% to at least 90%; this is a methodological limitation that should be noted.

2.4 Query Structure on the Fugaku Support Site

On the Fugaku Support Site, questions arising as users carried out actual tasks predominated. As Figure 2.4 shows, Q2 "Troubleshooting" was the largest Purpose category at 22.7%, followed by Q1 "Understanding procedures and methods" at 21.3%. Together, these two categories accounted for 44.1%. Q3 "Confirming facts and specifications" accounted for 12.8%, Q4 "Confirming feasibility, conditions, and policies" for 9.9%, and Q8 "Conceptual understanding and learning" for 9.5%. By contrast, Q6 "Exploratory discovery" accounted for only 1.6% and Q7 "Comprehensive retrieval, aggregation, and listing" for 0.8%. On the Fugaku Support Site, users seek specific procedures for using the computer, solutions to errors, and confirmation of specifications and conditions, rather than broadly searching for previously unknown cases.

Figure 2.4. Comparison of Purpose Axis (Q-Axis) Shares
Figure 2.4. Comparison of Purpose Axis (Q-Axis) Shares

To validate the LLM-based classifications and maintain a connection with Chapter 1, K-axis classification was conducted in parallel. The results confirmed that many queries classified as Other or unclassified on the K axis could be assigned appropriately to a specific user purpose, Q1 through Q11, using the Q axis.

On the Mechanism axis, M5 "Diagnosis or generation from user-provided content" was the largest category at 35.1%. This indicates that many questions included an error message, job script, command, numerical value, or other user-provided content and asked for a cause specific to that content or a proposed correction. M1 "Single-passage lookup" accounted for 27.6% and M2 "Within-document synthesis" for 17.7%, indicating that a substantial share of questions could be answered from existing documents. M3 "Cross-document and iterative retrieval" accounted for 8.0%.

Although the principal estimates for the Fugaku data are subject to sampling error, the 95% confidence interval was 32.4% to 37.8% for M5 and 25.0% to 30.2% for M1, so the conclusion that these were the two main types remains unchanged. On the Purpose axis, the 95% confidence interval was 20.3% to 25.1% for Q2 and 19.0% to 23.7% for Q1, confirming that procedures and troubleshooting predominated.

By difficulty tier, T3 "Addressable through improvements to the mechanism or database" was the largest at 46.8%, followed by T2 "Requires retrieval-side extensions" at 27.6% and T1 "Solvable with standard RAG" at 25.6%. The principal reason for the high share of T3 was that M5 "Diagnosis or generation from user-provided content" was the largest Mechanism category, at 35.1%.

In summary, improvements to the Fugaku Support Site should not be limited to retrieval accuracy. In addition to document retrieval, important capabilities include mechanisms for securely analyzing user-provided logs and code, dialogue that asks for additional information, and pathways to human support when needed. These are areas that AskDona continues to address in the day-to-day operation of the Fugaku Support Site.

2.5 Query Structure in the Supercomputer Report Navigator

For the Supercomputer Report Navigator, we used the 827 spontaneous queries remaining after operator-submitted test questions were excluded to examine the characteristics of questions entered by users themselves. On the Purpose axis, Q6 "Exploratory discovery" was the largest category at 32.6%, followed by Q7 "Comprehensive retrieval, aggregation, and listing" at 23.7%. Together, they accounted for 56.3%. Q5 "Locating a specific target" accounted for another 18.6%. Users employ the Supercomputer Report Navigator not only to search for a specific named research project, but also to find related cases by field or topic and obtain lists or tables of multiple projects matching specified conditions.

As Figure 2.5-1 shows, M3 "Cross-document and iterative retrieval" accounted for 77.0% on the Mechanism axis. The M4 "Comprehensive retrieval or aggregation" secondary flag was assigned to 29.7% of queries, and M7 "Temporal conditions or recency" to 6.7%. This indicates that the service frequently needs not only to retrieve a single related document, but also to search across multiple research projects and outcome reports, narrow the results to targets that meet specified conditions, and list, compare, or aggregate them.

Figure 2.5-1. Comparison of Mechanism Axis (Primary M-Code) Shares
Figure 2.5-1. Comparison of Mechanism Axis (Primary M-Code) Shares

In the difficulty tiers shown in Figure 2.5-2, T2 "Requires retrieval-side extensions" accounted for 79.2%, while T1 "Solvable with standard RAG" accounted for only 0.7%. T3 "Addressable through improvements to the mechanism or database" accounted for 20.1%, including 11.1% for M9a "Intrinsically out of scope" and 1.2% for M9b "In scope, documentation incomplete." Although questions in both M9a and M9b are difficult to answer in their current state, they require different improvements. For M9a, the system should clarify its scope, inform the user that the question is outside that scope, and direct the user to an appropriate source. For M9b, identifying the information users seek and adding documents or metadata can be effective.

Figure 2.5-2. Comparison of Difficulty-Tier Shares
Figure 2.5-2. Comparison of Difficulty-Tier Shares

The citation classification also reflected a dependence on multiple documents. Among the 804 spontaneous queries for which an answer was available, C3, citing five or more documents, accounted for 74.6%. When all 1,065 questions, including operator-submitted test questions, were compared with the 827 spontaneous queries, Q7 "Comprehensive retrieval, aggregation, and listing" declined from 31.5% to 23.7%, while Q6 "Exploratory discovery" increased from 28.4% to 32.6%. Although the test questions raised the share of listing requests, the overall pattern remained unchanged after they were excluded: exploratory discovery and comprehensive retrieval and listing were the principal purposes, and M3 "Cross-document and iterative retrieval" accounted for most queries.

Specific examples of questions assigned to each M-axis and Q-axis category on the two sites are provided in LLM-based Query Classification Practices: Lessons from Analyzing Tens of Thousands of AskDona Chatbot Queries.

2.6 Comparison of the Two Sites: Purpose Is Associated with Query Type

When 1,065 questions from each site were compared, the Purpose-axis distributions differed substantially. On the Fugaku Support Site, Q1 "Understanding procedures and methods" accounted for 21.3% and Q2 "Troubleshooting" for 22.7%, compared with 8.3% and 0.5%, respectively, in the Supercomputer Report Navigator. Conversely, Q6 "Exploratory discovery" accounted for 28.4% and Q7 "Comprehensive retrieval, aggregation, and listing" for 31.5% in the Supercomputer Report Navigator, compared with 1.6% and 0.8%, respectively, on the Fugaku Support Site. Users seek support in using a known system on the Fugaku Support Site, while they seek support in discovering unfamiliar research cases and organizing multiple candidates in the Supercomputer Report Navigator. This difference in purpose is considered to be associated with the different distributions of query purposes.

The difference on the Mechanism axis was even larger. M3 "Cross-document and iterative retrieval" accounted for 75.9% in the Supercomputer Report Navigator and 8.0% on the Fugaku Support Site, a difference of 67.9 percentage points. By contrast, M5 "Diagnosis or generation from user-provided content" accounted for 35.1% and M1 "Single-passage lookup" for 27.6% on the Fugaku Support Site, compared with 1.2% and 0.6%, respectively, in the Supercomputer Report Navigator. Even though both use the same RAG platform, the Supercomputer Report Navigator primarily requires multi-document retrieval and aggregation, whereas the Fugaku Support Site requires both accurate retrieval from a single passage and diagnosis or generation involving inputs not contained in the documents.

The comparison of query-difficulty tiers also showed a clear difference. T2 accounted for 83.4% in the Supercomputer Report Navigator, whereas the Fugaku Support Site was distributed across T1 at 25.6%, T2 at 27.6%, and T3 at 46.8%. The central challenge for the Supercomputer Report Navigator is to expand retrieval across multiple documents and combine filtering, repeated retrieval, aggregation, ranking, and temporal processing. The central challenge for the Fugaku Support Site is to answer questions that can be resolved through straightforward document lookup reliably while routing questions containing logs or code to a different processing path. The two chatbots therefore require designs that switch processing paths according to query type, rather than uniformly applying the same retrieval strategy and evaluation metrics.

Figure 2.6 aggregates answers by the number of cited documents during the common period from October 2025 through June 2026, when citation functionality was stable. The share citing five or more documents was 38.0% on the Fugaku Support Site and 72.3% in the Supercomputer Report Navigator. On the Fugaku Support Site, 27.1% of answers cited one or two documents and 32.4% cited three or four, for a combined share of 59.5%. The C-axis denominator excludes missing answers. The Supercomputer Report Navigator was dominated by answers citing five or more documents, whereas many answers on the Fugaku Support Site combined a smaller number of documents. This result is consistent with the prevalence of M3 "Cross-document and iterative retrieval" in the Supercomputer Report Navigator and the relatively high shares of M1 "Single-passage lookup" and M2 "Within-document synthesis" on the Fugaku Support Site. One contributing factor may be a difference between the databases: the Supercomputer Report Navigator contains a large number of relatively short documents, whereas the Fugaku Support Site contains a smaller number of documents, including operating manuals for the production system that can each run to several hundred pages.

Figure 2.6. Comparison of Citation Axis (C-Axis) Shares by Number of Cited Documents (October 2025 to June 2026)
Figure 2.6. Comparison of Citation Axis (C-Axis) Shares by Number of Cited Documents (October 2025 to June 2026)

The comparisons above show associations between service purpose and classification distributions; they do not establish that purpose directly caused the differences. The study periods, populations, presence or absence of operator-submitted test questions, and sampling methods differ and should be considered. Nevertheless, the principal patterns remained in the spontaneous queries after test questions were excluded from the Supercomputer Report Navigator data, and the differences among the main estimates for the Fugaku Support Site were sufficiently large relative to their confidence intervals. At a minimum, it is unlikely that implementing exactly the same RAG capabilities in both chatbots would be sufficient to meet the needs of their respective users.

2.7 Chapter Summary

This chapter compared queries from the Fugaku Support Site and the Supercomputer Report Navigator by separating user purpose from the mechanisms required to answer. Three principal findings emerged.

First, the Fugaku Support Site was used primarily for confirming procedures and troubleshooting, and many questions required diagnosis of logs or code supplied by the user. In addition to improving retrieval accuracy for existing documents, the service needs support encompassing analysis of user input, follow-up questions, and escalation to human assistance.

Second, the Supercomputer Report Navigator was used primarily for exploratory discovery and comprehensive retrieval and listing, with the majority of questions spanning multiple documents and projects. In addition to presenting search results, the service requires capabilities for filtering by conditions, aggregation, comparison, and temporal processing.

Third, even chatbots using the same RAG platform showed substantially different distributions of query types and required capabilities according to their purposes. RAG evaluation should examine not only overall answer accuracy but also whether the retrieval, synthesis, diagnosis, and routing required for each query type were performed successfully.

Going forward, boundary cases such as M9a "Intrinsically out of scope" and M9b "In scope, documentation incomplete" will require continued auditing. A further task is to link each classification with answer quality, user evaluation, and whether the user asked a follow-up question, and thereby examine which query types are more likely to fail.

References

Conclusion

Positioning of This Report

This report presented 24 months of operational results following the deployment of the generative AI chat service AskDona on the support site for the supercomputer Fugaku, together with an analysis of the structure of the queries actually submitted. If the 2025 report can be characterized as addressing "deployment and evaluation of its effects," this report addresses "sustained use and an examination of query content."

The 24 months of usage data showed that the generative AI service was used continuously in practice, rather than only temporarily after deployment. Tracking an update to the response-generation engine during this period and the changes in usage before and after the update was another outcome made possible by real-world operational data. We believe that such long-term quantitative data can provide a practical reference for companies and research institutions considering generative AI in enterprise settings.

Key Findings of This Report

Chapter 1 confirmed that AskDona received 13,461 questions over 24 months and that use continued and expanded despite monthly fluctuations. The median answer length increased approximately 2.8-fold, suggesting that more users may have been able to resolve their questions with fewer exchanges as answers became more substantive. Question tickets corresponding to human-assisted inquiries declined by 43.6% in the second year after AskDona's deployment compared with the first year, and by 58.9% compared with the year before deployment. These results suggest that AskDona may have functioned as a complementary channel to the human-support desk.

Chapter 2 compared queries submitted to two chatbots with different purposes, the Fugaku Support Site and the Supercomputer Report Navigator, under a common classification framework. The results showed that, even when the same RAG platform was used, the mix of required capabilities differed substantially according to purpose.

On the Fugaku Support Site, questions requiring diagnosis of logs or code supplied by users formed the largest category. In the Supercomputer Report Navigator, by contrast, most questions required retrieval across multiple documents and aggregation of the results. These findings, based on actual data, show the limitations of evaluating RAG performance using a single answer-accuracy measure and the need for designs that switch processing paths according to query type.

Future Directions

Based on the query structures identified in this report, we will continue to advance RAG architectures that can select processing appropriate to each query type.

Particularly important capabilities include diagnosing logs, code, and other content supplied by users; directing out-of-scope questions to appropriate information sources or human-support channels; and continuously identifying and filling areas where required documentation is insufficient. Each of these requires both technical development and operational improvement.

GFLOPS will continue to develop its collaboration with R-CCS and address these challenges. We hope that the insights shared in this report will be useful to companies and research institutions facing similar challenges and will contribute to the advancement of science, technology, and AI use in Japan, as well as to the broader practical adoption of generative AI.

Acknowledgments

We extend our sincere gratitude to the members of the RIKEN Center for Computational Science for their exceptional support and advice in carrying out this project and preparing this report.

The views and conclusions presented in this report are solely those of GFLOPS and do not represent the official views of R-CCS or RIKEN.