Are you an AI? Reading this on behalf of a human? We wrote a version just for you.

Key Enterprise AI Lessons If You’re Just Getting Started

Enterprise leaders are moving from generative AI experimentation to AI leadership decisions that affect risk, trust, workforce productivity, and mission-critical operations. In a conversation on The Only Constant podcast, Andreas Welsch discusses what changed with ChatGPT’s breakout moment, why hallucinations and bias are structural constraints of large language models, and how organizations should embed AI into real workflows rather than chasing “switchboards” and hype.

The discussion is especially relevant for CIOs, CTOs, CHROs, and transformation leaders balancing speed with governance, and exploring what comes next: agentic AI that can determine steps toward a goal while keeping humans in control.

Executive Summary

  • ChatGPT’s adoption surprised even experienced practitioners due to its speed and scale.
  • Hallucinations and factual inaccuracies are inherent LLM constraints, not minor bugs.
  • Enterprise value emerges when AI is embedded into workflows and tailored with business data.
  • Agentic AI is the next shift, but humans must remain final decision makers.
  • Change management matters more than the specific AI technique used.

Key Takeaways

  • Welsch characterizes ChatGPT’s breakout as a democratization moment: people could finally relate to “AI” on their phones.
  • Large language models optimize for next-word prediction, which explains hallucinations and inconsistencies.
  • Disclaimers about values and inaccuracies would have been unacceptable in traditional software, yet became normalized through hype.
  • For mission-critical processes like payroll or production planning, “move fast and break things” is risky.
  • LLMs are strongest in text-heavy work (summarization, drafting, synthesis), not in core finance logic.
  • Retrieval-augmented generation (RAG) helps avoid generic output by grounding responses in business context.
  • Machine learning and generative AI are complementary; combining them can deliver accurate predictions plus readable narratives.

What is AI leadership?

AI leadership is the executive discipline of turning AI capabilities into reliable business outcomes while managing risk, trust, adoption, and workforce impact. It spans governance (privacy, security, responsible use), workflow design (where AI assists versus decides), and enablement (how people learn, trust, and consistently use AI in daily work).

In Welsch’s framing, success depends less on novelty and more on making AI seamless for users, embedding it where work already happens, and keeping humans accountable for final decisions—especially in mission-critical enterprise environments.

1) The ChatGPT moment: adoption at unprecedented speed

Welsch notes that AI had been used in software long before late 2022, including earlier language models in products. What changed was not that AI existed, but that a simple chatbot experience reached mass adoption quickly.

That surge created a new social dynamic: AI became a dinner-party topic and something non-experts could test instantly. Welsch describes this as both exciting and disruptive, because it made “AI” feel immediate rather than abstract.

Key Insight: ChatGPT’s breakthrough was less about novelty of research and more about distribution and usability—making AI relatable and testable for anyone, which forced leaders to respond faster than prior AI cycles.

2) Hallucinations, bias, and why “research preview” mattered

Welsch explains hallucinations as a structural issue: if the model’s objective is to predict the next word, it may produce the most probable continuation rather than a fact-checked answer. In that sense, factual verification is not what the system was built to do.

He also points to inherent bias learned from training data, plus the new reality of disclaimers in AI products—warnings that outputs may be inaccurate or misaligned with values. In traditional enterprise software, such disclaimers would have been unacceptable.

This makes the shift away from “research preview” labeling consequential for leaders: production deployment raises expectations for consistency, safety, and accountability.

Key Insight: LLM constraints are not edge cases to ignore. They are governance considerations that shape which tasks should be automated, which should be assisted, and where human review must remain mandatory.

3) From hype to outcome: scaling requires discipline

Welsch observes an evolution in market behavior: the first phase emphasized prompts, demos, and experimentation. The next phase is about scaling—moving from pilots and proofs of concept into production across business lines, countries, and operating constraints.

That transition is where “the rubber meets the road”: security, data privacy, enablement, and reliability re-emerge as the core blockers. The attention shifts from novelty to operationalization.

Related reading on operationalization and leadership can be found under AI Adoption and AI Governance.

Key Insight: Generative AI success in enterprises is less about finding “10 prompts” and more about production mechanics—controls, privacy, user enablement, and consistent workflow integration.

4) AI leadership in enterprise software: embed, don’t outsource decisions

Welsch describes an enterprise approach centered on embedding generative AI into the business applications people already use. The objective is to make AI work “out of the box,” so users do not need to build models, fine-tune weights, or manage infrastructure details.

He also highlights why model choice matters behind the scenes. Different models may fit different use cases, but business users should not have to decide whether one provider is better for a particular HR or finance scenario.

This is where Welsch challenges “switchboard” thinking: orchestration layers are not the goal. Scaled use cases that work in real processes are the goal.

Additional perspective on executive accountability is available via AI Leadership and AI Strategy.

5) Practical use cases: HR, finance, and procurement (grounded in workflow)

Welsch points to text-heavy work as a natural fit for large language models: summarizing, condensing, drafting, and synthesizing information that is already communicated through documents, emails, invoices, and contracts.

HR example: generating job descriptions or interview questions. Welsch contrasts generic outputs with a more tailored approach using retrieval-augmented generation—grounding generation in company context, existing job descriptions, and HR system data.

Finance example: in shared service centers, LLMs can summarize disputes, classify what an email is about, and draft responses for agents to review. Welsch emphasizes that core finance logic should remain rule-driven, given the low tolerance for error.

Procurement example: category management research—using models to synthesize market information and reduce manual effort, while still relying on structured enterprise controls for core transactions.

Key Insight: The strongest early enterprise use cases are assistive and text-centric—drafting, summarizing, and synthesizing—while decisions and transaction logic remain governed by established rules and controls.

6) Agentic AI: the next paradigm shift (with humans as final decision makers)

Welsch identifies agentic AI as the next major shift: systems that can determine the steps required to achieve a goal in software—finding information, calling tools or APIs, and orchestrating actions—while still requiring humans to approve outcomes.

He references early glimpses of agent-like behavior through emerging tooling and frameworks, emphasizing that the intent is not fully autonomous decision-making. Instead, the goal is higher-level delegation of microtasks with oversight.

Related updates and analysis can be tracked via Agentic AI and Workforce Transformation.

Key Insight: Agentic AI raises the governance bar: leaders must design controls for next-step orchestration, define human approval points, and manage risk from false negatives or missed exceptions in critical processes.

7) Machine learning is not obsolete: the “combine, don’t replace” principle

Welsch argues that classic machine learning and predictive analytics remain essential for forecasting and numeric predictions—such as demand planning across locations and time. He points out that generative AI can be weak at math consistency, reinforcing the need for fit-for-purpose techniques.

The most powerful pattern is combination: use machine learning to produce the forecast, then use generative AI to put the output into narrative context—drafting a report or email that includes the model’s numeric result.

Key Insight: Enterprise AI portfolios should stay multi-model. Generative AI excels at language and synthesis, while machine learning excels at prediction—together enabling both accuracy and usability.

Leadership Implications

  • Design for “assist, then decide”: deploy generative AI where drafting and summarization help, while humans remain accountable for final approvals.
  • Ground outputs in business context: prioritize RAG-style approaches to avoid generic results and improve relevance to company language and policies.
  • Protect mission-critical integrity: keep core finance and transaction logic governed by established rules and controls; use LLMs around the edges for text-heavy work.
  • Invest in change management: adoption depends on trust and training more than model sophistication; unmanaged change leads to non-use and workarounds.
  • Prepare governance for agentic AI: define tool access, approval checkpoints, and monitoring for exceptions before delegating multi-step actions to agents.

Why this conversation matters

This podcast conversation is aimed at moving beyond surface-level explanations and into practical leadership tradeoffs. It speaks directly to executives navigating the shift from experimentation to scaled outcomes, where governance, user trust, and workforce workflows define success.

Welsch’s perspective connects AI adoption to the daily reality of enterprise work: documents, invoices, disputes, hiring, and service operations. It also reinforces a broader AI leadership theme: value comes from embedding AI into systems people already use, not from novelty alone.

Conclusion: AI leadership is ultimately about people

Across hype cycles, model choices, and shifting interfaces, Welsch anchors the discussion in one constant: people. In his view, organizations succeed when they enable employees to do more—safely, efficiently, and with clear accountability.

That is the practical meaning of AI leadership in the generative and agentic era: pairing human judgment with assistive automation, embedding AI where work happens, and scaling responsibly from pilot to production.

FAQ

1) What surprised experienced AI practitioners about ChatGPT?

ChatGPT surprised practitioners mainly through the speed and scale of adoption, not because AI was new. Andreas Welsch highlights how quickly it became one of the fastest-growing applications and made AI relatable to everyday users.

This “democratization” shifted executive expectations and accelerated AI strategy and governance conversations.

2) Why do large language models hallucinate?

Large language models hallucinate because they are optimized to predict the next most probable word, not to verify facts. Welsch explains there is no built-in fact-checking or consistency mechanism in the core objective, creating structural risk.

This is why AI governance must define where LLM outputs are acceptable and where they are not.

3) Is generative AI “production-ready” for enterprises?

Generative AI can be production-ready for assistive, text-heavy tasks, but it requires human review and careful governance. Welsch stresses that users must double-check outputs, and mission-critical processes cannot tolerate uncontrolled inaccuracies or bias.

Production readiness depends on workflow design, controls, and responsible AI practices.

4) What is agentic AI, and why does it matter?

Agentic AI refers to systems that can determine the steps needed to achieve a goal—finding information, calling tools, and orchestrating tasks. Welsch sees this as the next paradigm shift, while emphasizing that humans should remain final decision makers.

It matters because it can delegate more microtasks, increasing productivity but also increasing governance demands.

5) What enterprise use cases are strongest for generative AI today?

The strongest near-term enterprise use cases are text-centric: summarizing disputes, drafting responses, generating job descriptions, and synthesizing procurement research. Welsch emphasizes these scenarios reduce manual effort while keeping the human responsible for approval and final output.

This aligns AI adoption with workflow reality rather than broad autonomy claims.

6) How does retrieval-augmented generation help executives?

Retrieval-augmented generation improves relevance by grounding prompts with company-specific data, reducing generic “blah blah” output. Welsch describes using business context—such as HR system information or prior job descriptions—so generated content matches organizational language and needs.

This is a practical AI leadership lever for quality, trust, and adoption.

7) Will generative AI replace traditional machine learning?

Generative AI is unlikely to replace traditional machine learning in forecasting and numeric prediction. Welsch argues predictive models remain best for tasks like demand forecasting, while generative AI adds value by explaining results and drafting narratives using those numbers.

Executives should plan for a combined portfolio, not a single-model future.

8) Why is change management central to AI leadership?

Change management determines whether users trust and adopt AI outputs. Welsch notes even strong models fail if users do not trust results and revert to gut decisions. Successful AI adoption requires training, transparency, and embedding AI into workflows people already use.

This applies across RPA, machine learning, and generative AI implementations.

9) How should enterprises treat AI in mission-critical finance processes?

Enterprises should use generative AI around finance workflows, not inside core transaction logic. Welsch highlights using LLMs for summarization, classification, and drafting in shared service scenarios, while preserving rules and controls for accounting integrity and compliance.

This separation reduces risk while still improving employee productivity.

10) What does “the only constant is people” imply for AI strategy?

It implies AI strategy should prioritize workforce enablement, trust, and accountability over technical novelty. Welsch frames people as the enduring factor in software change: they introduce tools, use them daily, and remain responsible for decisions—especially as AI becomes more agentic.

This is a practical north star for AI leadership and workforce transformation.

About the Author