AI Agents Need Business Context to Run Enterprise Software Reliably
SAP chief executive Christian Klein argues that AI will change how enterprise software is used, not make systems such as SAP’s obsolete: agents still need reliable business data, process knowledge and controls to handle consequential work. In an interview with Big Technology, he says those requirements make accuracy and fragmented company systems key limits on adoption, even as he expects agents to take on more tasks. Klein also expects jobs to shift, with companies needing to reskill workers and change how they organize work.

Reliability depends on the task—and on the company’s data
For Christian Klein, the obstacle to putting AI into business operations is not simply whether a model can produce a useful answer. It is whether an agent can meet the accuracy a particular task requires, work from reliable data, and follow the rules governing what it may access and do.
Klein described AI’s progress to SAP employees as a climb from the cloud mountain to an AI mountain, with obstacles still to overcome. Generative AI made it easy to experiment with summarizing emails and documents, he said. Customers now want to know what that capability can do for the business. For mission-critical operations, fluent output is not enough: the agent must understand the process, handle the underlying data, and operate within governance requirements.
The accuracy threshold varies with the consequences of error. Klein gave financial closing as an example: he said an agent might understand the business at 93% accuracy, but that would not be enough if the numbers were wrong when results were released and audited. Separately, he said SAP has a rule that agents need to reach 95% accuracy or better to be worth shipping. The examples point to different thresholds for different tasks, rather than one standard for every business process.
Klein’s example of a customer briefing shows the other end of the range. He uses SAP’s Joule assistant to prepare for customer visits. If a briefing is 90% or 95% accurate, he said, it can still be useful; a slightly outdated earnings figure does not carry the same consequences as an error in financial results. The time saved preparing briefings can outweigh the need for perfect accuracy.
Klein expects the technology to improve quickly, but his forecast was qualified by task. He said reaching the required level for some applications within months was “absolutely possible.” He gave supply-chain optimization as an example of a further application that could take eight or nine months. SAP would keep a human in the loop, he said, while allowing customers to decide how autonomous an agent should be.
That forecast sits alongside a constraint Klein described at a customer visit: a financial-closing assistant was 93% accurate, while the customer had roughly 100 finance systems, some outside SAP. The issue was not necessarily the agent alone. The company’s data was scattered across systems, and data quality and silos made it harder to establish where financial information sat and how to reconcile it. If the system cannot reliably connect the records needed to answer a question, improving the model does not by itself resolve the problem.
The tension is central to Klein’s account of near-term progress. He expects agents to become trustworthy for some business tasks within months, while also describing customers whose fragmented systems hold those agents back. Better models may help clean up data, he said, but businesses still need connections among their systems and a way to establish what the records mean.
SAP is working to make that cleanup easier. Klein said its data platform partners with providers including Databricks, Snowflake, and BigQuery, and that AI can help match data across systems such as SAP, Salesforce, and Workday. The aim is to let agents use a semantic layer that relates the records, rather than requiring data teams to repeatedly reconcile finance, HR, payroll, and employee information. Klein said this work could reduce the need for armies of data scientists to perform monthly matching and cleanup.
That underlying context is part of the reliability problem, not an optional feature added after a model is built. An agent that detects a machine problem from sensor data, for example, also needs to relate that signal to maintenance orders, spare parts, and workers who can perform the repair. Without those links, it may identify an issue but fail to take the next useful step. Governance matters as well: Klein said agents working with public-sector customers need to know which data they may access, share, and store, including requirements such as FedRAMP.
Automation shifts work, but Klein does not expect the workload to disappear
Klein expects AI to speed up business execution, but says organizations will have to change their own operating rhythms to keep pace. SAP can no longer plan its agents separately from its employees, he said; workforce planning and virtual-agent planning need to be considered together. Planning cycles in areas such as finance, supply chain, and staffing will also have to get shorter as the technology moves faster.
He sees a change in customer attitudes, too. In earlier software transformations, businesses might migrate to a new system without changing how they worked, leaving employees uncertain about the point of the transition. Klein said AI has made customers more open to changing processes because they worry that failing to adapt could put their companies at risk. In his account, the pressure to change is not only coming from the technology; customers are increasingly willing to revisit how work is organized.
That change could remove substantial amounts of routine work. Klein named finance and HR controls, compliance checks, workflow approvals, and other tasks that occupy people full time. AI could also help with planning, pricing, and preparing business updates. The intended shift, he said, is toward more time for work that requires judgment: discussing strategy, deciding whether to change course, or preparing to present financial results.
There is a tension in that promise. Kantrowitz asked whether automating routine tasks could leave people with only intense, high-cognitive-load work—and make them more tired even if they were more productive. He described his own experience of using AI for invoicing and administrative tasks while spending more of his time on interview preparation and other demanding work. His question was not whether administrative automation was useful, but whether removing the pauses between tasks could make the remaining work exhausting.
Klein said he had not heard customers or end users complain of greater fatigue. He argued that people generally welcome getting rid of approvals and compliance checks so they can focus on what matters. But he also said people’s jobs will change dramatically and organizations will need to manage the transition. Automation, in his account, changes the type of work people do; whether that makes the work more exhausting depends in part on how people structure their days.
Klein described the change in his own work as well: customer briefings and other preparation are increasingly automated, leaving him more time to focus on content and discussions with his team about strategy, course corrections, and leadership decisions. He did not expect the change to make work more boring or exhausting. He did, however, acknowledge that employees regularly ask whether their jobs are secure and what restructuring might mean for them.
That concern matters alongside Klein’s claim that workload is not falling. SAP is operating at full speed, he said, and employees are not running out of work in development, finance, HR, or other parts of the company. AI may make employees more productive without reducing the amount of work the company expects them to do. Kantrowitz suggested that companies could use the productivity boost to pursue more of their roadmaps rather than simply do the same work with fewer people. Klein agreed that jobs and the workforce mix would change, but did not promise that every existing role would remain.
Klein told employees he expects SAP to have a different mix of job profiles within 12 months: fewer people in some roles and more hiring for skills such as data science and full-stack development. The challenge, as he described it, is to reskill employees where possible while also bringing in people with new skills. New hires need to work alongside experienced colleagues who understand the business and the software beneath the AI. That combination matters because technical skills alone do not replace domain knowledge.
Reskilling requires a different way of working, not just new tools
Klein pointed to SAP’s transition from on-premise software to cloud development as an example of reskilling in practice. Cloud development required a different way of working: code had to be tested automatically, shipped into production, and kept stable by the teams responsible for it. The change was not just a matter of learning new tools; it also changed who was responsible for the software once it was released.
He cited the rollout of Joule at work, which he said was being used by 80,000 SAP employees.
At first, some employees resisted, asking what the tool would mean for their jobs and whether they could trust its output. Training and coaching helped them see where it could save time and how to correct its work when needed. Klein said its use then spread across legal, HR, finance, and sales.
For Klein, the example illustrates why introducing a tool is not the same as changing how work gets done. Employees need to understand how to use it, where its output may need correction, and what they can do with the time it frees. The transition also depends on whether the new work is one people want to take on.
That question applies to product teams as well. Klein said SAP product managers had spent decades responding to customer requests for additional features. With AI, the expectation is changing: product managers should spend less time working only from inside the company and more time with customers, rethinking how supply chains, inventory optimization, and payroll could work. Instead of adding another feature to an established process, they may be asked to reconsider the process itself.
Klein said many employees find that prospect exciting because it gives them a chance to redesign business operations. But he did not suggest everyone will want to make the move. Some people may prefer the kind of work they already know, and he said not every employee will make the transition. Reskilling is possible, in his account, but not automatic: the company has to bring people along, and the new work may call for a different mindset as well as different skills.
AI moves software’s value upward without removing the need for business context
The fear that customers would use AI to recreate enterprise software rested on a real change in development productivity, Klein said. When the concern emerged, he asked SAP product managers which products could be reproduced through “vibe coding.” Their answer was that small tools with little process or data context might be replicable, but core systems were different.
An enterprise resource planning system, Klein said, has millions of data fields and large numbers of relationships among them. It also encodes industry processes, local requirements, and compliance obligations. He said SAP spends hundreds of millions each year keeping its software compliant. Reproducing an interface or a narrow feature is not the same as recreating the interconnected knowledge needed to run finance, payroll, or supply-chain operations.
That argument does not mean SAP can rely on its existing software unchanged. Klein said value is moving upward from the system of record into the agent layer. SAP needs its AI to be strong enough to defend the role of the underlying software.
Kantrowitz pressed Klein on whether increasingly capable models might eventually absorb the domain knowledge that makes those systems difficult to reproduce. Klein’s competitive argument was that SAP’s role is not only to store business data. Its platform encodes the industry and process knowledge, semantic relationships, and governance that agents need to act within an organization. A standalone language model, he argued, cannot run a warehouse or close a company’s books without those elements.
Klein described SAP’s AI foundation as a layer alongside the language model, with business data, process knowledge, knowledge graphs, and semantic structures. In an asset-maintenance example, an agent might use sensor signals to detect a machine problem, then connect that information to maintenance orders, spare parts, and workers who can perform the repair. That kind of process-specific connection is what SAP says its platform contributes.
The same distinction applies to rules for action. An agent handling a financial close needs to understand relevant tax requirements, Klein said; it cannot simply take action because a model has produced a plausible answer. Klein’s case for SAP’s continuing role was that the models are powerful and improving, but companies still need a platform that connects them to the way a business operates.
Klein expects companies to use multiple models rather than rely on one supplier. SAP makes models available on its platform and connects them to business context. He said he was not especially worried about a model provider withholding its latest model because there are many competing options, including open-weight models, and no single model solves every problem. That competition also lets SAP select a model for a task based on its performance and cost.
SAP allows outside agents in, but says Joule offers more than a connection
The question of who controls the AI interface is also a question of how agents reach enterprise data. SAP’s preferred interface is Joule, where users can work with models from providers such as Anthropic and Mistral while drawing on SAP’s business context. Klein said users could switch models for different tasks, including using Mistral where data sovereignty is important.
Third-party agents will also be able to access SAP systems. Klein said SAP would allow that through its API gateway and agent gateway, including for providers such as OpenAI and Anthropic. He did not describe an all-or-nothing policy in which outside agents are barred. His distinction was between access through a controlled gateway and using Joule, where SAP says it can supply richer semantic context, process knowledge, and governance.
A direct API or MCP connection can reach data, Klein said, but it does not by itself provide the same ontology or business context available on SAP’s AI platform. Nor does the connection automatically provide the same controls over who may see particular information. For example, when a finance user asks for an analysis through Joule, SAP can apply permissions to determine who is allowed to see the report.
Klein’s position is that SAP offers third-party agents a route into its systems while arguing that its own interface is the more capable route for many SAP users. An outside agent might access order data for a marketing campaign through the gateway. But a finance or HR employee working directly with SAP, Klein expects, will have reason to use Joule because it connects the model to business context and SAP’s governance.
Kantrowitz compared the choice to platforms that either allow outside assistants to act within their systems or require customers to use the platform’s own interface. Klein said SAP’s model includes both options. He expects Joule to be the main route, but not the only one.
Model selection makes token costs part of product design
The cost of running agents is becoming part of the product design, Klein said. SAP began building with leading frontier models, but has found that less expensive open-source models can perform well enough for tasks such as generating a headcount or financial report. The appropriate model depends on the accuracy the business needs, the model’s cost, and the consequences of an error.
SAP tests agents against different models and switches when another option offers a better balance between cost and performance. Klein said agents in development had already seen several model changes. Model selection is an ongoing product-management responsibility, not a decision made once at launch. He said SAP can test a new model and switch within a week or two, though the change may require adjustments to APIs and MCP servers.
Some work still calls for frontier models. Klein said SAP continues to use them for much of its coding, among other tasks. But he also pointed to a tabular AI model SAP had acquired, designed for predictions from structured data. In one retailer example, the model worked across SAP and non-SAP data and achieved the same accuracy for that customer as a team of ten data scientists, without a data scientist first building data pipelines or matching the records. Klein used the example to argue that business AI will involve different kinds of models, not just a choice between frontier and standard language models.
Rising productivity does not justify rising costs without limit. Klein said SAP’s overall token spending had increased, but a 20% productivity gain would not be worthwhile if costs rose 30%.
The company introduced token budgets by job profile roughly three or four months before the interview. Employees in finance, HR, product, and development received limits, with a process to request more for an exceptional task or project.
Klein said SAP did not impose the same limits on the top 1% of users working on model training and agent accuracy, because he considered that work essential. For other employees, he said, the budgets were intended to be high enough to support normal use without encouraging experimentation that had no corresponding productivity benefit. He described the policy as a compromise: employees could use AI in their jobs, and managers could approve more tokens for a specific task, while the company kept spending under control.
Klein said switching models could reduce costs by a factor of ten, making model choice more powerful than limits alone.
The underlying principle is that not every task needs the best available model. For some financial processes, SAP wants very high accuracy; for a customer briefing or headcount report, a lower level may still save enough preparation time to be worthwhile. A model has to fit both the task and the economics of delivering it.
More capable agents raise the stakes for security and regulation
Klein said SAP is testing multi-agent systems and evaluating how agents collaborate. He described model changes as technically manageable, but cybersecurity requires a different level of vigilance. Kantrowitz raised the risk of agents being directed to exploit vulnerabilities in systems holding sensitive business records. Klein agreed that cybersecurity must be at the top of the agenda for technology companies.
SAP is using AI to detect vulnerabilities earlier, test company and product defenses, and improve patching, Klein said. He confirmed that cybersecurity spending was increasing, but described the increase as reasonable: investment is going up as the company uses AI, while AI can also make finding vulnerabilities and patching systems more productive.
On European regulation, Klein said he believed regulators had good intentions, but argued that they should regulate AI’s effects on business and society rather than the technology itself. He said overlapping requirements can make it difficult for companies to understand how data may be used in AI research and customer applications. He singled out the AI Act and Data Act, rather than GDPR alone, as sources of additional layers and gray areas.
Klein said the European Commission was collecting simplification and deregulation measures under an “omnibus” effort, and he was confident that some rollback would follow. For startups in particular, he argued, regulatory complexity can make it harder to get a business going. His proposed distinction was between regulating AI’s impact on society and regulating the technology as such—a balance he said Europe would need to strike if it wanted to remain competitive.

