Prikazani su postovi s oznakom corporategovernance. Prikaži sve postove
Prikazani su postovi s oznakom corporategovernance. Prikaži sve postove

ponedjeljak, 10. kolovoza 2026.

AI Agents and Accountability

AI Agents and Accountability

Autor: Nermin Sefić

Agentic AI systems - AI that can independently execute multi-step tasks, make decisions within a defined scope, and interact with other systems without constant human supervision - represent one of the most significant capability shifts in…

Autonomous AI agents deliver real value, but also a new category of risk. A complete framework for oversight technology, organizational culture, and governance accountability.

Agentic AI systems - AI that can independently execute multi-step tasks, make decisions within a defined scope, and interact with other systems without constant human supervision - represent one of the most significant capability shifts in enterprise technology, and also one of the most consequential in terms of accountability and risk.

Delegating responsibility to a subordinate employee, with defined limits of authority and mandatory escalation for decisions above a certain threshold, is a well-established concept in organizational management. Agentic AI autonomy, in that sense, isn't conceptually new - it applies the same fundamental principle of bounded delegation that organizations have applied to human employees for decades.

What makes agentic autonomy qualitatively different lies in scale and speed. A human employee makes a limited number of decisions per day, each subject to natural, human constraints on speed and throughput. An AI agent can make thousands of decisions in the same window, meaning even a small error rate, negligible at the level of a single decision, can accumulate into a significant aggregate problem before human oversight has any chance to notice and intervene.

This difference in scale requires a different approach to oversight than the one applicable to human employees. Instead of relying exclusively on periodic, retrospective review - sufficient for human decisions made more slowly and in smaller volume - agentic systems require continuous, automated monitoring capable of detecting error patterns quickly, before accumulated damage grows significantly.

Organizations that approach agentic implementation seriously must invest in technical infrastructure specifically designed for overseeing autonomous systems, rather than relying on infrastructure originally built for traditional, human-led processes. This infrastructure typically includes several key components.

First, comprehensive logging of every decision the agent makes, including not just the final action but the context and reasoning that led to it, to whatever extent that's technically extractable from the system. This logging becomes critical when a decision is later reviewed, whether internally or in the context of an external audit or dispute.

Second, automated anomaly-detection systems capable of recognizing when agent behavior deviates from expected patterns, even when an individual decision doesn't explicitly cross a threshold that would require human approval. This early anomaly-detection capability distinguishes a robust oversight system from a simple, static rule set the agent follows without wider context.

Third, the ability to intervene quickly and granularly - not just a full agent shutdown, which can be an overreaction for minor issues, but the ability to narrow the agent's autonomy for specific decision categories while a larger problem is investigated, preserving functionality for parts of the system that aren't in question.

Technical infrastructure, however well designed, addresses only part of the challenge of agentic implementation. An equally significant, though less discussed, challenge lies in the cultural adaptation of human teams who work alongside, or oversee, autonomous agents.

Employees whose work shifts from directly executing tasks toward supervising and managing agents that execute those tasks undergo a significant role change that requires a different skill set - less operational execution, more critical judgment, the ability to quickly recognize when agent behavior deviates from expectations, and the confidence to intervene when necessary instead of passively accepting an agent decision simply because it appears authoritative.

Organizations that don't address this cultural transition explicitly, assuming employees will naturally adapt to the new dynamic, risk two opposite, equally problematic outcomes - excessive trust in agent decisions without sufficient critical review, or the opposite, excessive suspicion that negates the value of automation through constant, unnecessary second-guessing of every agent decision.

A legal framework specifically designed for liability questions tied to autonomous AI systems remains, in most jurisdictions, at an early stage of development. Existing liability concepts - developed for a context in which decisions were traditionally made by humans, or by software executing explicit, predefined instructions - don't always map cleanly onto a scenario where an agent makes decisions within a broader, but not fully specified, scope of autonomy.

Organizations implementing agentic systems today operate in this legally uncertain space, which increases the value of internal policies that clearly define the scope of autonomy and mandatory human involvement - not just as operational good practice, but as documentation that may become relevant if legal liability is ever disputed.

For organizations planning to implement agentic AI systems, a useful practical checklist before launch includes several key questions that deserve a clear answer before an agent gains access to production systems or real business decisions.

Is the scope of the agent's autonomy explicitly documented, including a clear boundary between decisions the agent may make independently and those requiring human approval? Does technical infrastructure exist for comprehensive logging of agent decisions, sufficient for later reconstruction and explanation if a decision is challenged? Is there a named individual or team with clear accountability for overseeing agent behavior - not diffuse, collective responsibility that in practice means no one is actually tracking it?

Is there a mechanism for fast, granular intervention if a problem is discovered, tested in advance rather than only theoretically available? Has the team working alongside the agent gone through explicit, not just implicit, preparation for the new working dynamic, including clear guidance on when it's appropriate to intervene versus accept an agent decision?

GNK ASG d.o.o. continues to monitor developments in this field as part of a broader approach to responsible artificial intelligence implementation within the GNK DINAMO Ltd. Group. The core conviction shaping this approach is that agentic autonomy, however technically impressive and operationally useful, delivers genuine, sustainable value to an organization only when paired with an equally deliberate, equally rigorous framework of accountability, oversight, and clearly defined organizational responsibility.

Organizations that approach agentic implementation with this balance - enthusiasm for technical capability tempered by serious consideration of the governance framework - position themselves to realize the real, long-term value of this technology, avoiding both excessive caution that negates the benefits of automation and excessive haste that exposes the organization to unnecessary, poorly understood risk.

Organizations that have successfully implemented agentic systems have rarely done so in one broad sweep across the entire organization. Instead, the successful pattern typically involves a carefully structured pilot program, limited to a single, clearly defined, low-risk function, before wider scaling.

Such a pilot typically begins with an agent assigned a narrow, well-understood scope of tasks - for example, automating routine classification of incoming customer inquiries by department, a task with a clear, easily verifiable correct answer and a relatively low cost of error if the agent gets an individual case wrong. This low cost of error lets the organization observe actual agent behavior in production, learn from inevitable early mistakes, and gradually build confidence in agent capability before expanding scope to higher-stakes, higher-consequence tasks.

During this pilot period, organizations that gain the most from the experience typically establish a formal, regular review process - weekly or biweekly meetings dedicated solely to analyzing agent behavior, including both successful cases and errors, gradually building a richer picture of the actual boundaries of agent reliability before deciding on broader scaling.

The natural tendency when evaluating agent performance is to focus exclusively on accuracy rate - what percentage of agent decisions were correct. While this metric remains important, organizations with longer experience with agentic systems often find that additional metrics offer equally valuable insight.

The distribution of errors deserves attention equal to their overall rate - an agent that errs evenly across all task categories carries a different risk profile than an agent whose errors concentrate in a specific, narrow category of scenarios, even if the overall accuracy rate is identical in both cases. This second situation, paradoxically, can be easier to manage - once the problematic category is identified, targeted intervention or additional human approval specific to that category addresses most of the risk without needing to broadly restrict agent autonomy.

Agent confidence in its own decisions, where technically available to measure, also offers a useful signal - an agent that expresses low confidence precisely in cases where it later errs demonstrates useful self-awareness that can be leveraged to automatically escalate low-confidence decisions toward human review, while an agent that shows equally high confidence regardless of actual accuracy offers less useful signal for this kind of automatic escalation.

Equally important as implementation criteria are predefined criteria for temporarily or permanently withdrawing an agent from production if a serious problem is discovered. Organizations that define these criteria only after a problem has already occurred, under the pressure of an active crisis, often make decisions under emotional pressure that in hindsight prove either overly reactive or insufficiently decisive.

Predefined, clear thresholds - for example, an error rate exceeding a certain percentage within a defined window, or any single error above a defined severity - remove some of that emotional component from the decision, enabling faster, clearer action once the criterion is met, without needing a fresh debate over whether the situation is serious enough to warrant intervention.

It's worth closing by recognizing that the current generation of agentic AI systems, however impressive compared to previous generations of automation, likely represents one point on a continuum that will keep evolving over the coming years. Organizations building governance and oversight frameworks for agentic autonomy today are building a capability that will likely require continuous adaptation as the technology itself keeps evolving.

This perspective suggests the value doesn't lie only in solving today's specific agent-governance questions, but in building a broader organizational capability - a culture of critical thinking about autonomy, technical oversight infrastructure, and clear decision processes - that will remain relevant and useful regardless of how agentic technology itself continues to develop in the years ahead.

Organizations that already have mature operational risk-management frameworks - developed over decades for traditional business processes - have a natural advantage when implementing agentic AI systems, because many underlying principles remain applicable, though they require adaptation to the specifics of agentic autonomy. Segregation of duties, for example, a principle long applied to human processes so that no single person holds complete control over a critical process without independent review, translates naturally to agentic systems too - an agent that initiates a transaction shouldn't be the same agent that approves it, even if it's technically possible to assign both roles to the same system.

A concrete, practical example of what a tiered escalation framework can look like in practice: an agent responsible for approving minor operating expenses might have full autonomy for amounts below a defined threshold, with a requirement for additional approval from a single human supervisor for amounts in a mid-range, and mandatory escalation to senior management, with detailed justification, for any amount above a defined upper limit.

Such a tiered structure, applied consistently across different categories of agent decisions, lets an organization realize significant operational efficiency for routine, low-risk decisions, while maintaining an appropriate level of human oversight precisely where it's most needed - for decisions with more significant potential consequences.

Agentic autonomy represents one of the most significant shifts in how organizations can structure work, offering real, measurable operational value through automating tasks that previously required continuous human engagement. That value, however, remains conditional on a deliberate, rigorous approach to the risk management this new capability inevitably brings - an approach that requires equal attention to oversight infrastructure, clarity of organizational accountability, and the cultural adaptation of teams working alongside agents every day.

There's a subtle but important difference between genuine oversight of an agentic system and micromanagement that in practice negates the value autonomy was supposed to deliver. Organizations that fail to recognize this difference risk implementing an agentic system that, despite its technical capacity for autonomous action, in practice requires such dense human review that the real operational benefit becomes marginal.

Healthy oversight focuses on tracking aggregate behavior patterns and escalating genuinely risky, unusual cases, allowing routine, low-risk decisions to pass without individual human review of every single instance. Micromanagement, by contrast, requires human approval for practically every agent decision, regardless of its actual risk level, thereby losing the very operational efficiency that was the primary motivation for introducing agentic autonomy in the first place.

Finding the right balance between these two approaches requires an iterative process - starting with more conservative, denser oversight during the early implementation period, and gradually, based on a documented, reliable track record of agent performance, expanding the scope of autonomy the agent enjoys without individual human review, while always maintaining clearly defined thresholds for cases that genuinely warrant escalation.

For board members who approve investments in agentic AI technology but don't necessarily have deep technical understanding of how it works, several questions deserve a place in the standard approval process, regardless of the specific technical implementation under consideration. Who will be named as specifically accountable for overseeing this agentic system after implementation, not just for its initial setup? What is the plan for gradually expanding the scope of autonomy, and what specific, measurable criteria must be met before each expansion step?

Is there a clear, pre-agreed process for the case where an agent makes an error causing real, measurable harm - not just a technical plan for pulling the agent, but also a communication plan toward affected parties, internal and external? These questions, while not requiring the board to have technical understanding of the AI technology itself, require the same level of governance seriousness the board would apply to any other significant operational change with potentially significant consequences.

GNK ASG d.o.o. recommends this kind of structured, governance-serious approach to any organization considering the implementation of agentic AI systems, recognizing that the technical sophistication of the technology itself doesn't diminish, but rather increases, the importance of an equally sophisticated governance framework surrounding it.

Beyond internal risk management, organizations implementing agentic systems that directly affect customers or external partners face an additional dimension of consideration - how transparently to communicate that an agentic system, rather than a human employee, is making or participating in making a specific decision affecting that external party.

This transparency isn't just a matter of regulatory compliance, though that's increasingly relevant as regulatory frameworks around the world begin explicitly addressing disclosure requirements for AI system use in decisions that significantly affect individuals. Transparency also shapes trust - customers and partners who discover, perhaps after the fact, that they believed they were communicating exclusively with a human decision-maker when an agentic system actually played a significant role, may experience that discovery as a breach of trust, regardless of whether the agent's decision itself was correct.

Organizations that proactively, clearly communicate the role of agentic systems in their processes, instead of leaving that role implicit or hidden, build longer-term, more robust trust with external parties, even if that transparency requires a more uncomfortable, more direct conversation in the short term about the boundaries and possible limitations of current technology.

Across every dimension covered in this analysis - technical infrastructure, organizational culture, legal context, performance metrics, and external transparency - one consistent theme runs through: agentic autonomy delivers genuine value only when paired with an equally serious, deliberate approach to the risk management that autonomy inevitably brings. Organizations that achieve this balance position themselves for the long-term, responsible use of one of the most significant technological capabilities available to the corporate world today.

GNK ASG d.o.o. remains committed to this deliberate, balanced approach across all future initiatives related to agentic AI technology within the GNK DINAMO Ltd. Group.

#AIAgents #ArtificialIntelligence #GNKASG #NerminSefic #GNKDINAMOLtd #SeficNermin


Cjelovit tekst i izvor: https://gnk-asg.hr/en/objave/ai-agents-accountability-implementation-framework/

Autor i urednička odgovornost: Nermin Sefić. Izdavač: GNK ASG d.o.o..

#GNKASG #GNKDINAMOLtd #NerminSefic #BusinessIntelligence #AIagents #artificialintelligence #corporategovernance

Halucinacije AI modela

Autor: Nermin Sefić Halucinacije velikih jezicnih modela dolaze u obliku identicnom tocnom odgovoru - visesloevita strategija ublazavanja ...