—For those stuck on the question, “So what should we actually adopt?”

This report is intended for executives, corporate IT teams, and digital transformation leaders who have begun evaluating generative AI tools but have been unable to reach a conclusion. It has been prepared from the perspective of GFLOPS, the company behind AskDona, as a commercially oriented document. Drawing on facts and data, it organizes the key considerations involved in tool selection and concludes by explaining which layer of the problem AskDona is designed to address.


1. Why Is Selecting a Generative AI Tool So Difficult?

Organizations everywhere are rushing to select and introduce generative AI tools. The discussion spans executive meetings, corporate IT teams, and business units, yet the more people talk, the harder it becomes to reach a conclusion.

If we compare generative AI with earlier devices, the closest parallels are PCs and smartphones. Those markets also had a proliferation of providers, making selection difficult. After Windows 95, corporate PC procurement remained an exhaustive exercise in comparing models, operating systems, and peripheral vendors until notebook shipments surpassed desktop shipments in 2000 (A History of Notebook Computers). Likewise, until around 2009—when the iPhone 3GS rapidly gained popularity in Japan—companies were still deciding among mobile phones, PHS devices, and smartphones (A Historical Timeline of Mobile Phones).

There is, however, one fundamental difference between generative AI and those earlier devices: it can receive a task and carry out the actions required to complete it.

With PCs, people still wrote the documents, sent the emails, and entered the numbers into Excel themselves. Even with smartphones, it was still the salesperson who decided whether or not to make a sales call. Humans were always the ones taking action.

Generative AI—and especially the “AI agents” that came into widespread use from 2025 onward—works differently. When a single instruction requires multiple actions, an agent can plan the necessary sequence itself and carry the process through to completion. Gartner described AI agents in 2025 as technology that acts autonomously toward a goal. A ChatGPT agent, for example, can be told to “submit my expenses” and then independently open the expense-management system, collect receipt data, and enter the required information into the reimbursement form (SoftBank: Looking Back on 2025, the First Year of AI Agents).

In other words, generative AI is not merely another device that extends human capabilities, as PCs and smartphones did. It enters the organization as a layer capable of reshaping the work itself.

That is why it becomes harder to know what to select the more carefully the issue is considered. Questions arise such as, “Is this an IT initiative?”, “Is it an HR initiative?”, and “Should the business unit own it?” More stakeholders become involved, opinions multiply, and the evaluation process grows longer. The discussion may eventually arrive at a surprisingly simplistic conclusion: “Let’s just use ChatGPT,” or “We already use Microsoft, so it has to be Copilot.”

This is not the fault of those making the decision. The discussion goes in circles because everything is grouped together under the single label of “generative AI tools,” without first separating the underlying issues.


2. Start with the Numbers: Where Japanese Companies Stand Today

Before debating the issue based on impressions, it is worth taking a clear-eyed look at where Japanese companies currently stand.

2.1 Overall Adoption Rates

In a joint survey conducted by The Yomiuri Shimbun and Teikoku Databank in March 2026, 34.6% of companies said they were using generative AI in their work. This was the first time the figure had exceeded 30% (BigGo Finance: Generative AI Adoption Tops 30% Among Japanese Firms).

According to the Japan Users Association of Information Systems (JUAS) in its Corporate IT Trends Survey 2025, 41.2% of surveyed companies—primarily listed enterprises—had introduced or were preparing to introduce language-based generative AI. That represented a 14.3-point increase from 26.9% in the previous year (JUAS Corporate IT Trends Survey 2025, Preliminary Results).

By industry, adoption was especially high in information-intensive and highly regulated sectors: 60.8% in social infrastructure and 54.4% in finance and insurance. One private-sector estimate suggests that Microsoft 365 Copilot has reached 94% penetration among Nikkei 225 companies (All About AI: Global AI Adoption).

Yet among small and medium-sized enterprises, company-wide adoption remains at only around 5%, or roughly 10% when limited departmental adoption is included (MONEYIZM: AI Adoption Among SMEs). Around half of SMEs say they have not established a clear policy. The result is a quiet but widening AI divide between large listed companies and smaller businesses.

2.2 International Comparison: Where Does Japan Rank?

The figures are even more sobering in international comparison.

CountryGenerative AI Usage Among Companies
China81.2%
United States68.8%
Japan27.0%

(Source: AX Inc.: Latest Generative AI Usage Rates for 2026)

In the 2026 global AI adoption rankings, the United States (over 85%), China (over 80%), and Singapore (78%) occupied the top three positions, while Japan only narrowly entered the bottom of the top ten (Second Talent: Top Countries with Highest AI Adoption Rates; All About AI).

OECD-related data also shows a wide gap at the individual level: 30.3% in Japan had used generative AI, compared with 68.8% in the United States and 81.2% in China. The individual usage gap is more than twofold (National Diet Library: OECD Survey on Individual Use; Nikkei: Individual Use of Generative AI).

2.3 The Impact Numbers Are Even More Concerning

A PwC survey conducted in spring 2025 across the United States, the United Kingdom, Germany, China, and Japan presents an even more sobering picture.

Across the four countries excluding Japan, an average of 86% of companies said they were achieving clear results from generative AI. In the United States, 51% said the results had exceeded expectations. In Japan, only 13% said the same.

Japan also ranked last among the five countries in formal integration into business processes—such as requiring AI use in contract review—with a rate of only 24%.

(Sources: PwC: Generative AI Survey 2025, Five-Country Comparison; ITmedia Enterprise; ASCII: Japanese Companies Use AI but Fail to Produce Results)

Japan therefore finds itself in the least cost-effective position: adoption is progressing, but the reported impact is only one-quarter to one-half of the levels seen elsewhere. It is hardly surprising that executives ask, “So how much value is this actually creating?” and receive no clear answer.


3. The Limits of the “Just Roll Out a Major LLM to Everyone” Approach

3.1 General-Purpose LLM Platforms Are the Easiest Place to Start

The first tools most organizations tend to consider are general-purpose LLM platforms from major vendors: OpenAI’s ChatGPT, Google’s Gemini, Anthropic’s Claude, and Microsoft’s Copilot.

In Japan, a number of companies began rolling out these general-purpose LLMs to all employees in 2025, almost as if they were issuing PCs. Microsoft 365 Copilot was particularly easy to approve as an extension of Office 365 and reportedly reached 94% penetration among Nikkei 225 companies.

3.2 The Reality of How They Are Used Is Bleak

According to a GXO report compiled as of October 2025, Microsoft 365 Copilot had exceeded 15 million paid licenses, but its workplace conversion rate—the proportion of paid licenses that translated into actual workplace use—was only 35.8%. In other words, roughly two out of every three employees with a paid license were making little or no use of it (GXO: The Reality Behind Microsoft Copilot’s 35.8% Adoption Rate).

An ITmedia survey from 2025 found that actual enterprise usage remained fragmented, with ChatGPT at 55.2% and Gemini at 15.7%. This points to a pattern in which companies initially introduce Copilot, fail to establish regular use, and then see employees move to other tools (ITmedia: Copilot Takes the Lead Among Mid-Sized and Large Enterprises).

Three recurring reasons are cited:

  • Concerns about data governance: organizations cannot work out how to connect their own data to an LLM safely.
  • Insufficient change-management budgets: too little budget reaches training and operations.
  • No internal AI champions: there is no one who can translate the technology into actual business workflows.

All three are organizational problems, not technical ones.

3.3 General-Purpose LLMs Are Fundamentally Tools for Individual Productivity

This is the central point. The LLMs offered by OpenAI, Google, Anthropic, and Microsoft are all primarily designed to improve individual productivity. The tasks involved—writing, coding, translation, summarization, and research—generally support the work of an individual.

The PC analogy is useful. Even if a company provides every employee with a high-performance PC, advanced functionality offers little value to someone who cannot use it effectively. Productivity may not improve as much as expected, while paper-based processes remain intact despite the availability of PCs.

Generative AI follows the same pattern. LLM responses are not perfect and can confidently present fabricated information as fact. Effective use requires fact-checking and appropriate prompting, so the results ultimately depend on each person’s ability to use an LLM effectively.

PwC’s finding that Japanese companies are failing to extract value ultimately leads back to this point: licenses have been distributed, but AI has not been embedded into business processes. That is why only 13% of Japanese companies report results that exceed expectations (EnterpriseZine: Three Measures Suggested by the PwC Survey).


4. McKinsey’s “Two Stacks”: Both Horizontal and Vertical Adoption Are Stalled

PwC’s finding that only 13% of Japanese companies are seeing better-than-expected results suggests that the issue is not simply a difference in AI literacy. There are bottlenecks in the structure of AI adoption itself.

McKinsey’s 2025 State of AI report frames this as a global pattern. It organizes enterprise generative AI adoption along two dimensions:

  • Horizontal: broad, company-wide deployment with relatively shallow use
  • Vertical: narrowly targeted, function-specific deployment with deeper use

McKinsey argues that both are stalled, each for different reasons.

Horizontal tools such as company-wide Copilot deployments and chatbots can spread quickly, but their benefits tend to be diffuse and difficult to measure. More transformative Vertical use cases, meanwhile, remain stuck at the pilot stage in roughly 90% of cases.

(McKinsey: The State of AI 2025; November 2025 PDF)

Deloitte’s State of AI in the Enterprise 2026, based on 3,235 executives and IT leaders across 24 countries, found that more than two-thirds of executives said that 30% or less of their experiments had reached production. Multiple independent surveys therefore confirm the same barrier: organizations struggle to move from proof of concept to production (Deloitte).

4.1 The Horizontal Barrier: “We Distributed It, but People Do Not Use It”

The Horizontal approach—providing Copilot, ChatGPT, or Gemini to all employees—can be rolled out quickly. Microsoft 365 Copilot reportedly reached 94% of Nikkei 225 companies.

Yet as discussed in Chapter 3, the workplace conversion rate is only 35.8%, meaning two-thirds of paid seats remain largely unused (GXO). Across Japanese companies, only 13% say that generative AI has produced better-than-expected results in PwC’s five-country comparison (PwC).

This is precisely what McKinsey means by “diffuse, hard-to-measure gains.” Distribution was fast. But formal integration into business processes remains at only 24% in Japan, the lowest of the five countries surveyed by PwC. The licenses are in place, but their value is not being realized.

4.2 The Vertical Barrier: “We Ran a PoC, but It Never Reached Production”

The Vertical approach develops deeper use cases for specific functions or departments. McKinsey reports that roughly 90% remain in pilot mode.

The business functions most frequently identified as leading areas of generative AI use are broadly consistent:

IT / knowledge management / marketing and sales / product and service development / service operations and customer support / software engineering

(McKinsey: How Organizations Are Rewiring to Capture Value)

Companies are testing deeper use cases in each of these areas. Typical examples include:

  • Internal help desks for IT and corporate services
    The same questions about expense reimbursement, account provisioning, PC problems, and administrative procedures bounce repeatedly among HR, general affairs, IT, and finance. Even when FAQs exist, employees often ask someone on Slack first.

  • First-line customer support
    Immediate answers based on product specifications and FAQs. AskDona’s support implementation for RIKEN’s Fugaku reduced one inquiry from as much as four hours to five seconds, a representative example of this category.

  • Software engineering
    Code generation and review, consultation of internal documentation, and incident investigation. McKinsey identifies this as one of the areas with the clearest cost reductions.

  • Marketing and sales
    First drafts of proposals and copy, retrieval of past cases and product specifications, and organization of competitive information. McKinsey identifies marketing and sales as one of the areas showing the clearest revenue impact.

  • Legal and compliance
    Reviewing approximately 50 checklist items for every contract, including NDA provisions, governing law, and liability caps, as well as vendor assessments and system risk assessments. These are typical structured evaluation tasks spanning McKinsey’s “Risk & Legal” and “Knowledge Management” categories.

  • Knowledge management for policies and manuals
    Providing access to employment rules, expense policies, information-security policies, and compliance rules; operating evaluation guidelines and grading definitions; and searching product and operating manuals.

Across these areas, companies such as Ricoh, NEC, and Brains Technology point to the same persistent problem: documents exist, but remain effectively unusable.

“The materials are there, but cannot be found when needed.”
“Employees have to ask the person who knows.”
“Keyword-based internal search cannot understand intent or tacit knowledge.”

(Ricoh; NEC; Brains Technology)

The underlying solution is already well known: allow an LLM to reference internal documents using RAG—Retrieval-Augmented Generation.

But even when a PoC reports 83% or 90% accuracy, production rollout often stalls because of incomplete documents (46%), retrieval quality (42%), access controls, document revision and maintenance, and the path through which employees actually use the system (Canon IT Solutions: How RAG Changed Internal Search).

When McKinsey says that 90% of Vertical use cases remain stuck in pilot mode, it is describing exactly this phenomenon: the move from a successful PoC to production is precisely where deployment breaks down.

4.3 Where Japanese Companies Stand: Broad but Shallow Horizontally, Stalled Vertically

Combining the findings from McKinsey, Deloitte, and PwC produces a clear picture of Japanese enterprise adoption.

LayerCurrent StateIndicators
Horizontal: company-wide Copilot/ChatGPT deploymentDistribution has progressed, but adoption and impact remain limited94% penetration among Nikkei 225 companies / 35.8% workplace conversion / 13% reporting better-than-expected results
Vertical: deep, function-specific RAG use casesPoCs are conducted, but most do not reach productionRoughly 90% remain pilots (McKinsey) / fewer than one-third of executives have scaled more than 30% of experiments (Deloitte)

Both layers are stalled at the same time, but for different reasons.

Simply rolling out a general-purpose LLM does not deepen the impact of Horizontal adoption.
At the same time, if every Vertical PoC rebuilds internal data connections, model selection, security requirements, and operational design from scratch, budgets and timelines run out before production is reached.

This is where AskDona has a role to play.


5. What Problems Does AskDona Solve?

AskDona, developed by GFLOPS, is positioned as a platform that brings together:

  • the Horizontal layer—a secure enterprise GPT experience provided to individual employees
  • the Vertical layer—deep, function-specific use cases

Rather than addressing McKinsey’s two stalled stacks through separate vendors and separate operating models, AskDona is designed to solve them through a single platform, a single governance model, and a single user interface.

5.1 Product Overview: Four Core Capabilities

At the core of AskDona is proprietary RAG technology developed and optimized through repeated joint validation with RIKEN from the company’s earliest stages.

The project began with former Google employees and developed through close collaboration with RIKEN to put next-generation RAG technology into practical use. That history forms the technical foundation of the current product (AskDona Official Website; GFLOPS Press Release).

AskDona is no longer limited to two product categories and now provides four core capabilities.

  • RAG Chat
    This core capability coordinates multiple AI agents to generate highly accurate answers from internal knowledge. Its source-reference feature allows users to preview the original files supporting an answer directly in the interface. PDF preview is supported, with Office-file support planned from September 2025 onward. Transparency into the supporting evidence—one of the most important ways to reduce hallucinations—is built into the user experience.

  • GPT Chat
    A general-purpose LLM assistant. In addition to GPT-4o, Claude Sonnet 4 was added on May 23, 2025, with unlimited questions. Frequently used prompts can be stored and reused as templates. This addresses the layer in which employees use AI for their individual work.

  • RAG × Deep Research (First in Japan)
    Five AI agents with different roles, including researchers and analysts, autonomously and recursively research internal data and generate structured reports. Rather than merely searching, the system carries out research, analysis, and summarization as one end-to-end process.

  • Batch Assessment—Professional Feature
    A function for automatically evaluating hundreds or thousands of items in a single batch. Use cases include compliance checks, vendor assessments, system risk assessments, and new-employee onboarding checklists. Assessments can be configured through an eight-step wizard, and results can be exported to Excel with assessment results, comments, and source citations.

AskDona is therefore no longer simply “a secure enterprise ChatGPT plus internal RAG.”

It now covers higher-level forms of knowledge use within a single platform:

Search with RAG Chat → research with Deep Research → evaluate and make judgments with Batch Assessment

5.2 AskDona’s Strengths in Numbers and Specifications

  • Outperforms major RAG frameworks by 16 to 28.9 percentage points on complex-question accuracy, demonstrated through joint validation with RIKEN.

  • Designed so that answer accuracy is less likely to decline even as data volume grows, a critical requirement when working with enterprise-scale document collections.

  • Uses semantic chunking, which divides content by meaning rather than character count, and metadata tagging, including titles, creation dates, and URLs, to improve both retrieval accuracy and source traceability.

  • Multi-LLM support enables organizations to use ChatGPT, Azure OpenAI, Gemini, and Claude for different workloads, avoiding lock-in to a single provider.

  • Supported file formats include PDF, Word, Excel, PowerPoint, CSV, TXT, HTML, Markdown, JSON, PNG, JPG, and public web pages.

  • Security includes ISO/IEC 27001:2022 certification (AT Press), TLS encryption in transit, AES-256 encryption at rest, storage of all data in data centers located in Japan, and a policy that input data is never used to train LLMs.

  • Access control includes the standard Administrator and Member roles, as well as custom roles defined by individual capabilities in accordance with the principle of least privilege.

  • IP restrictions support IPv4, IPv6, and CIDR notation, enabling controls based on internal networks and office locations.

  • Audit logs record all important operations—including additions, deletions, updates, and downloads—in chronological order to support compliance audits.

5.3 Representative Use Cases

(a) Enterprise Customer Support and Knowledge Access: RIKEN Fugaku

RIKEN introduced AskDona for user support for the Fugaku supercomputer. A first-line response that had previously taken as long as four hours was reduced to five seconds (GFLOPS / RIKEN Press Release; PR TIMES).

This is a rare Japanese example of a Vertical AI use case that moved beyond the pilot stage and scaled into production—despite McKinsey’s finding that roughly 90% of Vertical use cases remain pilots.

(b) Legal: Batch Contract Review with Batch Assessment

Legal teams often review approximately 50 checklist items for every new contract, including NDA clauses, governing law, liability caps, and anti-social-force provisions.

With Batch Assessment, a contract PDF is uploaded to the RAG database and assessed against checklist items defined in Excel. AI agents then classify each item as clause present / clause absent / requires review.

Legal professionals can focus their attention on the “clause absent” and “requires review” items, reducing both review time and the risk of oversight.

(c) Human Resources: Onboarding and Periodic Assessments with Batch Assessment

Department-specific onboarding checklists can be filtered and applied to new employees. The same assessment can also be rerun quarterly to observe changes over time.

This enables organizations to standardize training and onboarding processes and track progress without allowing the process to become dependent on individual staff members.

(d) Company-Wide Access to Internal Policies, Manuals, and FAQs with RAG Chat

Employment rules, expense policies, information-security policies, compliance rules, and product manuals can be consolidated in a RAG database.

The Insights page analyzes user feedback and questions the AI could not answer, allowing the knowledge base to be improved continuously. A Shared Sessions feature allows teams to reuse useful conversations and prompts.

(e) RAG × Deep Research: Delegating Research and Analysis Work

Multiple AI agents recursively investigate internal data and generate research reports.

Use cases include market analysis, identifying patterns across past projects, and producing first drafts of product-comparison materials—work that previously required analysts to spend many hours manually gathering and organizing information.

Use cases (b), (c), and (e) are important differentiators because they show that AskDona extends beyond search into research, analysis, and evaluation—the area of decision support.

Many of the Vertical use cases that McKinsey identifies as stuck in pilot mode break down at exactly this point: organizations can build a RAG PoC, but they do not reach structured evaluation, judgment, or deeper investigation.

AskDona incorporates these capabilities from the outset.

5.4 A Message for Companies That Said, “Let’s Start with Copilot for Now”

Based on the discussion so far, the message for companies whose evaluation has stalled is straightforward.

Company-wide deployment of a general-purpose LLM—the Horizontal approach—can improve individual productivity. But as the 35.8% workplace conversion rate shows, more than half of the licenses remain dormant if the tool is simply distributed.

Deep, function-specific use cases—the Vertical approach—may produce promising PoCs, but McKinsey reports that roughly 90% do not reach production. The reason is that organizations rebuild internal data connections, operating processes, security controls, and evaluation logic from scratch for every PoC.

AskDona is designed to address the weakness highlighted by PwC’s finding that only 24% of Japanese companies have embedded generative AI into formal business processes by providing:

  • Horizontal capabilities through GPT Chat
  • Vertical capabilities through RAG Chat, Deep Research, and Batch Assessment
  • all on the same platform, under the same governance model, and through the same user interface

The RIKEN Fugaku case—from four hours to five seconds—shows that a Vertical use case can reach production.

Deep Research and Batch Assessment show that AI can move beyond search and into decision support.


6. Three Ways to Reframe the Discussion When Tool Selection Has Stalled

For organizations unable to reach a conclusion, the discussion can be reframed around three axes. These are consistent with the findings from McKinsey, Deloitte, and PwC.

First, separate Horizontal deployment from Vertical specialization.

Company-wide distribution of general-purpose LLMs such as ChatGPT, Copilot, Gemini, and Claude belongs to the Horizontal layer. Function-specific RAG use cases such as help desks, customer support, and internal-policy retrieval belong to the Vertical layer.

As McKinsey shows, the reasons for failure are different:

  • Horizontal adoption produces benefits that are diffuse and difficult to measure.
  • Vertical adoption leaves approximately 90% of use cases stuck in pilot mode.

Mixing both into the same meeting almost inevitably turns the discussion into an abstract debate. They should be managed using different metrics and assigned to different owners.

Second, measure license distribution separately from business-process integration.

The number of seats distributed is no longer sufficient as a KPI. PwC’s findings show that the percentage of workflows formally incorporating AI—24% in Japan, less than half the U.S. level—is much more closely linked to realized impact.

McKinsey’s observation that Horizontal deployment spreads rapidly but produces diffuse benefits, and Deloitte’s finding that 30% or less of experiments reach production, ultimately come down to the same issue: organizations need a KPI for formal integration into business processes.

Third, decide in advance where AI may take action.

As generative AI becomes more agentic, the old assumption that “a human presses the button” begins to break down.

Organizations should define the following before selecting a tool:

  • access rights to internal data
  • authority to call external APIs
  • the workflow for final human approval

Defining these boundaries before selecting a tool can make the selection process considerably faster.


7. Conclusion

Generative AI tool selection stalls not because the people making the decision are at fault, but because three distinct dimensions are being mixed together:

  • Horizontal distribution vs. Vertical specialization
  • License distribution vs. business-process integration
  • Tool selection vs. action authority

Japanese companies now find themselves in an unfavorable position: adoption has exceeded 30%, yet realized impact remains only one-quarter to one-half of global levels.

McKinsey’s 2025 report describes the structure clearly:

  • the Horizontal layer spreads quickly, but its benefits remain diffuse
  • roughly 90% of Vertical use cases remain pilots

Deloitte similarly found that more than two-thirds of executives had moved 30% or less of their experiments into production.

Japan is therefore stalled at both layers simultaneously.

AskDona integrates:

  • the Horizontal layer, through GPT Chat as a secure enterprise AI assistant distributed to individuals
  • the Vertical layer, through RAG Chat, Deep Research, and Batch Assessment for function-specific knowledge work

The product is designed to eliminate both traps—“distributed but unused” and “successful PoC but no production rollout”—through one platform.

The RIKEN Fugaku case, in which response time fell from four hours to five seconds, is one of the few Japanese references demonstrating that a Vertical use case can scale into production.

For organizations whose discussion has stalled, the first question should be:

“Are we stuck on the Horizontal layer, the Vertical layer, or both?”

When the answer is “both,” operating two separate foundations through two separate vendors is likely to create duplication in governance, operations, and cost. A single platform capable of supporting both Horizontal and Vertical adoption offers a more effective model.

AskDona is one of the few options currently available that is designed to move the entire discussion forward at once.


References and Sources

International Comparisons

Global Research: McKinsey and Deloitte

Copilot and LLM Adoption Challenges

AI Agents and RAG

AskDona and GFLOPS

History of PC and Smartphone Adoption


This draft is version 3. It preserves the original argument—how generative AI differs from PCs and smartphones, the limits of general-purpose LLMs, the challenges of function-specific adoption, and AskDona’s positioning—while restructuring Chapter 4 onward around McKinsey’s 2025 Horizontal/Vertical framework. Chapter 5 has also been rebuilt using the detailed AskDona report dated May 8, 2026, incorporating the platform’s four core capabilities—RAG Chat, GPT Chat, RAG × Deep Research, and Batch Assessment—and its specific use cases across RIKEN support, legal, HR, knowledge access, and research. Further claims, figures, or case studies can be added, removed, or replaced after review.