Human Judgment Is Contextual—and Sometimes Wrong

People do not apply rules in a vacuum. They interpret rules in light of context, purpose, experience, and perceived risk. This flexibility enables people to respond to circumstances that were not fully anticipated when a rule was written. At the same time, it can produce inconsistent decisions, particularly when people act on incomplete information or assumptions that cannot later be verified.

Consider a pedestrian waiting at a red light on a quiet side street. The road appears empty, the pedestrian knows that it is rarely busy, and crossing immediately may seem harmless. That conclusion is based on context and experience. However, it may still be wrong. A vehicle may be outside the pedestrian’s field of view, its speed may be misjudged, or an unusual event may occur precisely when the pedestrian assumes that the road will remain empty.

The same conflict appears when a person’s destination is directly across an unsignalized section of road but a designated crossing is farther ahead. Convenience and personal risk assessment may suggest crossing immediately, while the formal safety system directs the person toward the designated crossing.

In Japan, the appropriate action at a red pedestrian signal is to wait. Where a crosswalk or signalized intersection is nearby, pedestrians are generally expected to use it. This is not merely a matter of formal compliance. According to Japan’s National Police Agency, 4,158 fatal collisions between motor vehicles and pedestrians occurred during the five years from 2021 through 2025. Approximately 70% involved pedestrians who were crossing the road. Of those crossing fatalities, approximately 60% occurred outside a crosswalk, and legal violations were identified in approximately 70% of those cases. [1]

Traffic laws and social norms differ across jurisdictions, but the underlying cognitive problem is universal: people everywhere make decisions from partial observation, prior experience, and perceived risk. Japan provides one concrete example; it does not define the boundaries of the issue.

The important lesson is not that breaking a rule is a sign of being human. It is that a decision can feel reasonable to the person making it while still being based on incomplete observation, overconfidence, or an incorrect assessment of risk. Human judgment is valuable, but it is not automatically correct.

Judgment Is Influenced by Culture and Experience, but Not Determined by Them

People’s interpretation of rules is influenced by education, professional experience, social norms, organizational practices, cultural context, and personal history. These factors shape tendencies; they do not determine how every individual will behave.

An observational study conducted at selected pedestrian crossings in Nagoya and Strasbourg found that 41.9% of the observed crossings in Strasbourg occurred against a red light, compared with 2.1% in Nagoya. The finding indicates that social norms and surrounding behavior can influence risk-taking and rule compliance. It does not prove that every Japanese person follows rules or that every French person ignores them. It was a comparison of particular locations and populations, not a measure of fixed national character.

Cross-cultural research has also distinguished relatively “tight” cultures, which have stronger social norms and lower tolerance for deviation, from relatively “loose” cultures, which permit a broader range of behavior. Neither is inherently superior. Strong norms can support coordination and consistency, while greater flexibility can support adaptation and individual expression. These are group-level patterns and should never be used to predict an individual’s competence or judgment solely from nationality or cultural background.

This diversity in human judgment is one of the sources of creativity, adaptability, and innovation. In high-stakes environments, however, the same diversity can also create a problem: two qualified people may review the same information and reach different conclusions, yet neither may be able to reconstruct exactly why their conclusion differed from the other.

The Challenge in Regulated Work

This issue becomes particularly important in areas connected to people’s health, safety, rights, and financial well-being, including healthcare, financial services, compliance, insurance, public safety, and parts of the criminal-justice system.

In these domains, decisions are rarely based on a single rule or a single document. A professional may need to consider laws, regulations, internal policies, technical manuals, historical cases, contractual conditions, customer information, risk indicators, and exceptions. Some sources may be incomplete, outdated, or inconsistent with one another.

Organizations have automated clearly defined calculations, routing processes, scoring models, and rule checks for many years. The central technological problem was therefore not that automation itself was impossible. The difficult problem was automating or supporting complex judgment while satisfying all of the following conditions at the same time:

  • reconciling contradictions across multiple sources;
  • interpreting context and legitimate exceptions;
  • identifying information that is missing;
  • recognizing when the available evidence is insufficient;
  • abstaining from a conclusion when confidence is inadequate;
  • allowing a third party to trace the evidence behind a recommendation;
  • testing whether outcomes are fair across relevant groups;
  • establishing who remains accountable for an incorrect decision; and
  • updating the system when regulations, policies, or source documents change.

This is where large language models combined with retrieval-augmented generation, or RAG, may add value beyond conventional search and fixed rule engines.

A properly designed RAG system can retrieve relevant passages from controlled organizational documents, compare information across multiple sources, synthesize the evidence, identify potential gaps, and produce a preliminary assessment with references. It can help turn scattered documents into a reviewable decision process.

It should not, however, be described as possessing the same understanding, accountability, or professional responsibility as a human expert. The objective is not to make the language model the final authority. The objective is to make evidence collection, policy checking, and preliminary analysis faster, more consistent, and easier to audit.

AI Should Support the Decision Process, Not Own the Decision

In regulated environments, the most important question is not simply whether AI can be introduced. It is how the system is governed, validated, monitored, and connected to accountable human decision-makers.

The EU AI Act requires high-risk AI systems to support risk management and testing, record relevant events through logs, disclose their capabilities and limitations, and provide effective human oversight. People responsible for oversight must be able to understand the system’s limitations, remain alert to automation bias, disregard or reverse an output, and interrupt the system when necessary. [2]

The World Health Organization similarly emphasizes human autonomy, safety, transparency, explainability, responsibility, accountability, inclusiveness, and equity in the design and use of AI for health. [3]

Financial regulation also illustrates why traceability matters. In the United States, Regulation B requires the principal reasons for an adverse credit decision to be specific and to accurately describe the factors that were actually considered or scored. A statement that a customer simply failed to meet an institution’s internal standard is insufficient. [4]

The 2026 revised U.S. interagency model-risk-management guidance emphasizes risk-based testing, validation, monitoring, governance, defined responsibilities, and controls for the models within its scope. Importantly, the revised guidance expressly excludes generative AI and agentic AI because those technologies are novel and rapidly evolving. This means that the guidance should not be presented as a complete governance framework for LLM applications; organizations need controls designed specifically for each generative-AI use case. [5]

GFLOPS’s position should therefore be expressed as follows:

AI should not indiscriminately replace accountable human judgment. It should make evidence gathering, source comparison, policy checking, missing-information detection, and preliminary assessment faster, more consistent, and more auditable. Uncertain, exceptional, or high-impact cases should be escalated to qualified human reviewers who retain the authority and responsibility for the final decision.

A Global Organizational Challenge with Different Local Pathways

Critical organizational knowledge often becomes concentrated in particular people. This is not unique to any one country.

Across industries and markets, experienced employees learn how to interpret policies, reconcile apparently conflicting sources, recognize when an exception may apply, identify which evidence matters, and decide when a case should be escalated. Written procedures may describe the standard rule without fully recording how experts handle ambiguity or borderline cases.

Japan illustrates one pathway by which organizational knowledge becomes person-dependent. Other countries reach the same structural problem through different pathways.

Japan: Long Tenure and Firm-Specific Expertise

Japan has historically developed company-specific expertise through long employee tenure, on-the-job training, and internal knowledge transfer. OECD data published in 2026 reported an average job tenure of 12.4 years in Japan, with 46.7% of workers having remained with the same employer for at least ten years. [6] At the same time, these employment practices are gradually changing as labor mobility increases and organizations face new skill requirements.

This model has enabled organizations to build deep firm-specific expertise. An experienced employee may understand not only what a manual says, but also which document takes priority, how a rule has historically been interpreted, what evidence is normally accepted, and when an exceptional case should be referred to another department or senior reviewer.

The relevant business problem is not simply that Japanese companies spend a large amount on training. The stronger problem statement is that critical organizational knowledge can become concentrated in particular people.

Common examples include:

  • knowledge being concentrated in a small number of experienced employees;
  • written procedures being insufficient to reproduce how experts handle exceptions;
  • tacit knowledge being lost when employees transfer or retire;
  • new employees spending substantial time searching for the correct information;
  • different employees giving different answers to the same question; and
  • the reasoning behind past decisions not being recorded in a reusable form.

Internal training has played an important role in Japanese companies, but the OECD notes that it often concentrates on firm-specific skills and is becoming less suitable for a labor market with greater mobility and rapidly emerging skill requirements. [6]

In Japan, this vulnerability may become especially visible when long-serving employees transfer or retire before their practical knowledge has been made explicit.

Other Markets: Mobility, Scale, and Organizational Fragmentation

Outside Japan, the same structural problem may emerge through different conditions.

In labor markets with higher employee mobility, knowledge may leave the organization more frequently through resignation, contracting, outsourcing, or changes in project teams. In rapidly growing companies, hiring and expansion may move faster than documentation and specialist development. Following a merger, acquisition, or restructuring, different teams may continue to use different definitions, systems, approval practices, and interpretations even after adopting a common group policy.

Multinational organizations face another version of the problem. Relevant knowledge may be divided among headquarters, regional offices, local subsidiaries, shared-service centers, and external advisers. A global policy may be written in one language, implemented through local procedures in another, and subject to different legal requirements in each jurisdiction.

The route differs, but the underlying organizational question is the same:

Can the organization reproduce a reliable decision process without depending on one individual’s memory, personal network, or undocumented experience?

A company may possess all the information required to make a decision while still being unable to use that information consistently. The relevant material may be distributed across documents, systems, departments, languages, and jurisdictions. A manual may state the general rule without recording the practical questions an expert asks, the evidence the expert prioritizes, or the conditions under which an exception should be escalated.

The business case for AI is therefore global. It is not simply reducing training expenditure. It is transforming dispersed and person-dependent knowledge into a process that is searchable, explainable, reviewable, transferable, and adaptable to local requirements.

A More Diverse, Mobile, and Global Workforce Requires More Explicit Organizational Knowledge

Countries and organizations face different workforce pressures. Japan provides a particularly visible example through population ageing, a shrinking labor force, and the gradual transformation of long-standing employment practices. Under a scenario in which fertility, net immigration, and employment rates remain constant, the OECD projects that Japan’s labor force would shrink by more than half by the end of the century. The OECD therefore recommends a combination of productivity growth and greater participation by women, older people, and foreign workers rather than reliance on any single solution. [6]

Japan had 2,571,037 foreign workers as of October 2025, the highest number recorded since the current reporting requirement began. [7]

Other markets may encounter the same organizational challenge through specialist shortages, higher turnover, international recruitment, remote work, rapid expansion, or cross-border operating models. The demographic and labor-market conditions are not identical, but each can weaken processes that depend on long-term personal familiarity with unwritten rules.

This development should not be described by claiming that immigrants or internationally recruited employees inherently make different or less reliable judgments. Such language risks connecting nationality or background with professional competence.

The organizational issue is different. As workforces become more diverse, mobile, distributed, and cross-functional, companies can no longer assume that every employee shares the same language, educational background, regulatory experience, institutional history, professional terminology, or unstated workplace conventions. These differences can exist between countries, but also between departments, professions, generations, acquired companies, and employees who joined the same organization at different times.

The weakness lies not in workforce diversity. It lies in an organization’s dependence on assumptions and practices that have never been made explicit.

Organizations therefore need to:

  • document decision criteria clearly;
  • identify the authoritative global and local sources for those criteria;
  • provide accessible and, where appropriate, multilingual materials;
  • distinguish formal legal or policy requirements from customary practices;
  • record how legitimate exceptions should be handled;
  • make the evidence behind a decision visible;
  • define when an employee or AI system must escalate a case; and
  • ensure that equivalent cases can be assessed according to equivalent standards while respecting local requirements.

A diverse workforce makes this need more visible, but the resulting improvements benefit everyone. Explicit criteria and traceable evidence reduce dependence on personal networks, institutional memory, lengthy tenure, and familiarity with unwritten organizational conventions.

The common global problem can therefore be stated as follows:

Important organizational decisions often depend on information and judgment that are fragmented, difficult to transfer, and insufficiently traceable. The local causes differ, but the need to make the decision process explicit is shared.

A Japanese Deployment Addressing a Global Knowledge Problem

GFLOPS has been exploring how RAG-based systems can support knowledge-intensive work through AskDona.

The deployment described below is based in Japan, but the underlying problem is widely shared: users must navigate extensive specialist documentation, relevant information may be distributed across multiple sources, and the complete question is not always apparent at the beginning of the interaction.

RIKEN Center for Computational Science states that it has introduced AskDona on the Fugaku support website. RIKEN describes AskDona as a generative-AI assistant based on RAG technology that generates responses by referring to Fugaku manuals and technical documentation, helping users address common questions about system usage. [8] Our own account of this work is published as AskDona for HPC.

This achievement should not be described by saying that an LLM has “understood Fugaku’s complicated manual” in a human-like sense. A more accurate statement is:

AskDona retrieves and synthesizes relevant information from Fugaku’s extensive manuals and technical documentation, helping users locate and interpret information required to resolve common and complex technical-support questions.

A GFLOPS-published operational report describes a comparative evaluation using 239 Fugaku-related specialist documents totaling more than 16,000 pages. Twenty-five composite questions, selected by R-CCS experts to reflect actual inquiry patterns, were evaluated by three R-CCS specialists. The evaluators reviewed answers independently, with system identities concealed and answer order randomized. Answers were scored on accuracy, completeness, and practical usefulness. [9]

In this evaluation, AskDona dona-rag-2.0 received a normalized composite score of 83 out of 100, while the four comparison systems averaged approximately 61. This result must not be reported as “83% accuracy.” It was a normalized composite score based on three human-evaluated dimensions: accuracy, completeness, and practical usefulness.

The appropriate description is:

In a 25-question domain-expert evaluation conducted with R-CCS specialists, AskDona dona-rag-2.0 received a normalized composite score of 83 out of 100 across accuracy, completeness, and practical usefulness, compared with an average of approximately 61 for the four comparison RAG systems.

The result is promising, but it should be interpreted as an operational, domain-expert evaluation rather than an independent third-party benchmark. The question set was limited to 25 questions, the questions were selected by R-CCS experts, and the evaluators were also R-CCS specialists. Further external and independently reproducible evaluation would be required before making broader performance claims.

The same care is required when describing the effect on support volume. According to the GFLOPS report, after the ticket-creation process was routed through AskDona in February 2025, question-category support tickets were approximately 40% lower year over year in the first quarter of 2025 and approximately 53% lower in the second quarter. April 2025 showed an approximately 61% year-over-year reduction.

The accurate wording is:

Following the introduction of AskDona and a change that routed ticket creation through the system, human-handled question tickets were substantially lower year over year during the reported period, including a reduction of approximately 61% in April 2025.

It would be too strong to state that AskDona alone caused the entire reduction, because the AI deployment and the support-workflow change occurred together. The figures demonstrate a meaningful operational association, but they do not isolate every causal factor.

The report also presents an individual inquiry for which a human-prepared answer had required approximately four hours, while AskDona generated a response in approximately five seconds. This is a useful illustration of the potential speed difference for an information-retrieval-intensive question, but it should be identified as a single case rather than an average response-time result.

What the Fugaku Deployment Demonstrates—and What It Does Not

The Fugaku deployment provides evidence that a source-grounded LLM system can support complex technical-information retrieval, cross-document synthesis, and user self-service in a demanding scientific-computing environment.

It does not, by itself, establish that the same system can autonomously make medical, financial, insurance, compliance, or criminal-justice decisions.

Those domains may involve direct effects on a person’s health, assets, opportunities, rights, or legal status. Applying the same architectural principles to such work would require additional controls, including:

  • domain-specific and independently reviewed test sets;
  • validation against real-world case distributions;
  • separate measurement of severe errors and minor errors;
  • testing of whether the system correctly abstains when information is insufficient;
  • fairness and disparate-impact evaluation where relevant;
  • source and policy version control;
  • access controls for personal and confidential information;
  • mandatory human review for defined categories of decisions;
  • monitoring of overrides, appeals, incidents, and repeated errors; and
  • clear assignment of final legal and professional accountability.

The most defensible conclusion is therefore:

The Fugaku deployment is a Japanese implementation addressing a globally shared knowledge problem. It demonstrates the potential of source-grounded AI to support knowledge-intensive technical work. Extending that approach to regulated or high-impact decisions in any market requires domain-specific and jurisdiction-specific validation, stronger governance, and qualified human accountability.

Reallocating Human Judgment, Not Eliminating It

GFLOPS does not need to argue that human judgment should disappear. Human judgment remains essential when values conflict, facts are incomplete, consequences are serious, or an exception cannot be resolved through existing policy.

The appropriate role for AI is to reduce the repetitive burden surrounding that judgment.

AI can help gather the relevant documents, compare requirements, identify missing information, expose contradictions, apply routine criteria, and prepare a traceable preliminary assessment. Human professionals can then devote more attention to exceptions, communication, empathy, creativity, ethical trade-offs, explanation, and final accountability.

Humanity is not defined by inconsistency or by the willingness to ignore rules. It is expressed through the ability to understand why rules exist, recognize when a situation requires further consideration, accept responsibility for consequences, and make value-sensitive decisions that cannot be reduced to mechanical processing.

The goal is therefore not to automate “human judgment” as a single undifferentiated activity. It is to separate the work that can be standardized and audited from the work that requires distinctly human responsibility.

GFLOPS aims to make organizational decision processes more evidence-based, consistent, and auditable by using AI to support information retrieval, criteria checking, missing-information detection, and preliminary assessment—while preserving human authority for uncertain, exceptional, and high-impact decisions.

This is not the replacement of human expertise. It is a way to preserve that expertise, distribute it more effectively, and focus it where it matters most.

References

  1. National Police Agency (Japan), Traffic fatality statistics, 2021–2025.
  2. European Union, Regulation (EU) 2024/1689 (Artificial Intelligence Act).
  3. World Health Organization, Ethics and Governance of Artificial Intelligence for Health.
  4. Consumer Financial Protection Bureau, Regulation B, 12 CFR Part 1002 (Equal Credit Opportunity Act).
  5. Board of Governors of the Federal Reserve System, Office of the Comptroller of the Currency, and Federal Deposit Insurance Corporation, Interagency Guidance on Model Risk Management, 2026 revision.
  6. OECD, Japan labour-market statistics and projections, 2026 (job tenure, employer-provided training, and labour-force projections).
  7. Ministry of Health, Labour and Welfare (Japan), Notification status of foreign workers’ employment, October 2025.
  8. RIKEN Center for Computational Science, Introduction of AskDona on the Fugaku support website.
  9. GFLOPS Co., Ltd., Fugaku × AskDona operational report. See the report.