Prikazani su postovi s oznakom artificialintelligence. Prikaži sve postove
Prikazani su postovi s oznakom artificialintelligence. Prikaži sve postove

ponedjeljak, 10. kolovoza 2026.

AI Agents and Accountability

AI Agents and Accountability

Autor: Nermin Sefić

Agentic AI systems - AI that can independently execute multi-step tasks, make decisions within a defined scope, and interact with other systems without constant human supervision - represent one of the most significant capability shifts in…

Autonomous AI agents deliver real value, but also a new category of risk. A complete framework for oversight technology, organizational culture, and governance accountability.

Agentic AI systems - AI that can independently execute multi-step tasks, make decisions within a defined scope, and interact with other systems without constant human supervision - represent one of the most significant capability shifts in enterprise technology, and also one of the most consequential in terms of accountability and risk.

Delegating responsibility to a subordinate employee, with defined limits of authority and mandatory escalation for decisions above a certain threshold, is a well-established concept in organizational management. Agentic AI autonomy, in that sense, isn't conceptually new - it applies the same fundamental principle of bounded delegation that organizations have applied to human employees for decades.

What makes agentic autonomy qualitatively different lies in scale and speed. A human employee makes a limited number of decisions per day, each subject to natural, human constraints on speed and throughput. An AI agent can make thousands of decisions in the same window, meaning even a small error rate, negligible at the level of a single decision, can accumulate into a significant aggregate problem before human oversight has any chance to notice and intervene.

This difference in scale requires a different approach to oversight than the one applicable to human employees. Instead of relying exclusively on periodic, retrospective review - sufficient for human decisions made more slowly and in smaller volume - agentic systems require continuous, automated monitoring capable of detecting error patterns quickly, before accumulated damage grows significantly.

Organizations that approach agentic implementation seriously must invest in technical infrastructure specifically designed for overseeing autonomous systems, rather than relying on infrastructure originally built for traditional, human-led processes. This infrastructure typically includes several key components.

First, comprehensive logging of every decision the agent makes, including not just the final action but the context and reasoning that led to it, to whatever extent that's technically extractable from the system. This logging becomes critical when a decision is later reviewed, whether internally or in the context of an external audit or dispute.

Second, automated anomaly-detection systems capable of recognizing when agent behavior deviates from expected patterns, even when an individual decision doesn't explicitly cross a threshold that would require human approval. This early anomaly-detection capability distinguishes a robust oversight system from a simple, static rule set the agent follows without wider context.

Third, the ability to intervene quickly and granularly - not just a full agent shutdown, which can be an overreaction for minor issues, but the ability to narrow the agent's autonomy for specific decision categories while a larger problem is investigated, preserving functionality for parts of the system that aren't in question.

Technical infrastructure, however well designed, addresses only part of the challenge of agentic implementation. An equally significant, though less discussed, challenge lies in the cultural adaptation of human teams who work alongside, or oversee, autonomous agents.

Employees whose work shifts from directly executing tasks toward supervising and managing agents that execute those tasks undergo a significant role change that requires a different skill set - less operational execution, more critical judgment, the ability to quickly recognize when agent behavior deviates from expectations, and the confidence to intervene when necessary instead of passively accepting an agent decision simply because it appears authoritative.

Organizations that don't address this cultural transition explicitly, assuming employees will naturally adapt to the new dynamic, risk two opposite, equally problematic outcomes - excessive trust in agent decisions without sufficient critical review, or the opposite, excessive suspicion that negates the value of automation through constant, unnecessary second-guessing of every agent decision.

A legal framework specifically designed for liability questions tied to autonomous AI systems remains, in most jurisdictions, at an early stage of development. Existing liability concepts - developed for a context in which decisions were traditionally made by humans, or by software executing explicit, predefined instructions - don't always map cleanly onto a scenario where an agent makes decisions within a broader, but not fully specified, scope of autonomy.

Organizations implementing agentic systems today operate in this legally uncertain space, which increases the value of internal policies that clearly define the scope of autonomy and mandatory human involvement - not just as operational good practice, but as documentation that may become relevant if legal liability is ever disputed.

For organizations planning to implement agentic AI systems, a useful practical checklist before launch includes several key questions that deserve a clear answer before an agent gains access to production systems or real business decisions.

Is the scope of the agent's autonomy explicitly documented, including a clear boundary between decisions the agent may make independently and those requiring human approval? Does technical infrastructure exist for comprehensive logging of agent decisions, sufficient for later reconstruction and explanation if a decision is challenged? Is there a named individual or team with clear accountability for overseeing agent behavior - not diffuse, collective responsibility that in practice means no one is actually tracking it?

Is there a mechanism for fast, granular intervention if a problem is discovered, tested in advance rather than only theoretically available? Has the team working alongside the agent gone through explicit, not just implicit, preparation for the new working dynamic, including clear guidance on when it's appropriate to intervene versus accept an agent decision?

GNK ASG d.o.o. continues to monitor developments in this field as part of a broader approach to responsible artificial intelligence implementation within the GNK DINAMO Ltd. Group. The core conviction shaping this approach is that agentic autonomy, however technically impressive and operationally useful, delivers genuine, sustainable value to an organization only when paired with an equally deliberate, equally rigorous framework of accountability, oversight, and clearly defined organizational responsibility.

Organizations that approach agentic implementation with this balance - enthusiasm for technical capability tempered by serious consideration of the governance framework - position themselves to realize the real, long-term value of this technology, avoiding both excessive caution that negates the benefits of automation and excessive haste that exposes the organization to unnecessary, poorly understood risk.

Organizations that have successfully implemented agentic systems have rarely done so in one broad sweep across the entire organization. Instead, the successful pattern typically involves a carefully structured pilot program, limited to a single, clearly defined, low-risk function, before wider scaling.

Such a pilot typically begins with an agent assigned a narrow, well-understood scope of tasks - for example, automating routine classification of incoming customer inquiries by department, a task with a clear, easily verifiable correct answer and a relatively low cost of error if the agent gets an individual case wrong. This low cost of error lets the organization observe actual agent behavior in production, learn from inevitable early mistakes, and gradually build confidence in agent capability before expanding scope to higher-stakes, higher-consequence tasks.

During this pilot period, organizations that gain the most from the experience typically establish a formal, regular review process - weekly or biweekly meetings dedicated solely to analyzing agent behavior, including both successful cases and errors, gradually building a richer picture of the actual boundaries of agent reliability before deciding on broader scaling.

The natural tendency when evaluating agent performance is to focus exclusively on accuracy rate - what percentage of agent decisions were correct. While this metric remains important, organizations with longer experience with agentic systems often find that additional metrics offer equally valuable insight.

The distribution of errors deserves attention equal to their overall rate - an agent that errs evenly across all task categories carries a different risk profile than an agent whose errors concentrate in a specific, narrow category of scenarios, even if the overall accuracy rate is identical in both cases. This second situation, paradoxically, can be easier to manage - once the problematic category is identified, targeted intervention or additional human approval specific to that category addresses most of the risk without needing to broadly restrict agent autonomy.

Agent confidence in its own decisions, where technically available to measure, also offers a useful signal - an agent that expresses low confidence precisely in cases where it later errs demonstrates useful self-awareness that can be leveraged to automatically escalate low-confidence decisions toward human review, while an agent that shows equally high confidence regardless of actual accuracy offers less useful signal for this kind of automatic escalation.

Equally important as implementation criteria are predefined criteria for temporarily or permanently withdrawing an agent from production if a serious problem is discovered. Organizations that define these criteria only after a problem has already occurred, under the pressure of an active crisis, often make decisions under emotional pressure that in hindsight prove either overly reactive or insufficiently decisive.

Predefined, clear thresholds - for example, an error rate exceeding a certain percentage within a defined window, or any single error above a defined severity - remove some of that emotional component from the decision, enabling faster, clearer action once the criterion is met, without needing a fresh debate over whether the situation is serious enough to warrant intervention.

It's worth closing by recognizing that the current generation of agentic AI systems, however impressive compared to previous generations of automation, likely represents one point on a continuum that will keep evolving over the coming years. Organizations building governance and oversight frameworks for agentic autonomy today are building a capability that will likely require continuous adaptation as the technology itself keeps evolving.

This perspective suggests the value doesn't lie only in solving today's specific agent-governance questions, but in building a broader organizational capability - a culture of critical thinking about autonomy, technical oversight infrastructure, and clear decision processes - that will remain relevant and useful regardless of how agentic technology itself continues to develop in the years ahead.

Organizations that already have mature operational risk-management frameworks - developed over decades for traditional business processes - have a natural advantage when implementing agentic AI systems, because many underlying principles remain applicable, though they require adaptation to the specifics of agentic autonomy. Segregation of duties, for example, a principle long applied to human processes so that no single person holds complete control over a critical process without independent review, translates naturally to agentic systems too - an agent that initiates a transaction shouldn't be the same agent that approves it, even if it's technically possible to assign both roles to the same system.

A concrete, practical example of what a tiered escalation framework can look like in practice: an agent responsible for approving minor operating expenses might have full autonomy for amounts below a defined threshold, with a requirement for additional approval from a single human supervisor for amounts in a mid-range, and mandatory escalation to senior management, with detailed justification, for any amount above a defined upper limit.

Such a tiered structure, applied consistently across different categories of agent decisions, lets an organization realize significant operational efficiency for routine, low-risk decisions, while maintaining an appropriate level of human oversight precisely where it's most needed - for decisions with more significant potential consequences.

Agentic autonomy represents one of the most significant shifts in how organizations can structure work, offering real, measurable operational value through automating tasks that previously required continuous human engagement. That value, however, remains conditional on a deliberate, rigorous approach to the risk management this new capability inevitably brings - an approach that requires equal attention to oversight infrastructure, clarity of organizational accountability, and the cultural adaptation of teams working alongside agents every day.

There's a subtle but important difference between genuine oversight of an agentic system and micromanagement that in practice negates the value autonomy was supposed to deliver. Organizations that fail to recognize this difference risk implementing an agentic system that, despite its technical capacity for autonomous action, in practice requires such dense human review that the real operational benefit becomes marginal.

Healthy oversight focuses on tracking aggregate behavior patterns and escalating genuinely risky, unusual cases, allowing routine, low-risk decisions to pass without individual human review of every single instance. Micromanagement, by contrast, requires human approval for practically every agent decision, regardless of its actual risk level, thereby losing the very operational efficiency that was the primary motivation for introducing agentic autonomy in the first place.

Finding the right balance between these two approaches requires an iterative process - starting with more conservative, denser oversight during the early implementation period, and gradually, based on a documented, reliable track record of agent performance, expanding the scope of autonomy the agent enjoys without individual human review, while always maintaining clearly defined thresholds for cases that genuinely warrant escalation.

For board members who approve investments in agentic AI technology but don't necessarily have deep technical understanding of how it works, several questions deserve a place in the standard approval process, regardless of the specific technical implementation under consideration. Who will be named as specifically accountable for overseeing this agentic system after implementation, not just for its initial setup? What is the plan for gradually expanding the scope of autonomy, and what specific, measurable criteria must be met before each expansion step?

Is there a clear, pre-agreed process for the case where an agent makes an error causing real, measurable harm - not just a technical plan for pulling the agent, but also a communication plan toward affected parties, internal and external? These questions, while not requiring the board to have technical understanding of the AI technology itself, require the same level of governance seriousness the board would apply to any other significant operational change with potentially significant consequences.

GNK ASG d.o.o. recommends this kind of structured, governance-serious approach to any organization considering the implementation of agentic AI systems, recognizing that the technical sophistication of the technology itself doesn't diminish, but rather increases, the importance of an equally sophisticated governance framework surrounding it.

Beyond internal risk management, organizations implementing agentic systems that directly affect customers or external partners face an additional dimension of consideration - how transparently to communicate that an agentic system, rather than a human employee, is making or participating in making a specific decision affecting that external party.

This transparency isn't just a matter of regulatory compliance, though that's increasingly relevant as regulatory frameworks around the world begin explicitly addressing disclosure requirements for AI system use in decisions that significantly affect individuals. Transparency also shapes trust - customers and partners who discover, perhaps after the fact, that they believed they were communicating exclusively with a human decision-maker when an agentic system actually played a significant role, may experience that discovery as a breach of trust, regardless of whether the agent's decision itself was correct.

Organizations that proactively, clearly communicate the role of agentic systems in their processes, instead of leaving that role implicit or hidden, build longer-term, more robust trust with external parties, even if that transparency requires a more uncomfortable, more direct conversation in the short term about the boundaries and possible limitations of current technology.

Across every dimension covered in this analysis - technical infrastructure, organizational culture, legal context, performance metrics, and external transparency - one consistent theme runs through: agentic autonomy delivers genuine value only when paired with an equally serious, deliberate approach to the risk management that autonomy inevitably brings. Organizations that achieve this balance position themselves for the long-term, responsible use of one of the most significant technological capabilities available to the corporate world today.

GNK ASG d.o.o. remains committed to this deliberate, balanced approach across all future initiatives related to agentic AI technology within the GNK DINAMO Ltd. Group.

#AIAgents #ArtificialIntelligence #GNKASG #NerminSefic #GNKDINAMOLtd #SeficNermin


Cjelovit tekst i izvor: https://gnk-asg.hr/en/objave/ai-agents-accountability-implementation-framework/

Autor i urednička odgovornost: Nermin Sefić. Izdavač: GNK ASG d.o.o..

#GNKASG #GNKDINAMOLtd #NerminSefic #BusinessIntelligence #AIagents #artificialintelligence #corporategovernance

AI in Credit Risk Assessment

AI in Credit Risk Assessment

Autor: Nermin Sefić

How AI is changing credit risk management - expanded signals, continuous monitoring, the explainability problem, and regulatory compliance. An in-depth analysis.

How AI expands available signals, enables continuous monitoring, and why model explainability becomes as important as its accuracy.

When discussing artificial intelligence in a corporate context, the conversation almost inevitably turns toward generative models, chatbots, and content automation. Considerably less attention goes to a question that could have greater long-term financial impact for most medium and large organisations: how AI is changing the very nature of credit and counterparty risk management.

Classical credit risk assessment models, built on decades of statistical practice, rely on a relatively small number of structured variables - payment history, debt-to-equity ratio, liquidity indicators, external agency credit ratings. These models work well for stable, well-documented business entities with long histories, but systematically underperform for new, fast-growing, or structurally unusual business partners - precisely the categories organisations increasingly deal with in a globalised, rapidly changing economic environment.

The problem is not the mathematics of traditional models itself, but the limited breadth of data they rely on. A business partner founded two years ago, growing quickly but lacking sufficiently long history to generate a reliable credit rating through traditional agencies, represents a blind spot for organisations relying exclusively on classical models - despite real, measurable signals about their health potentially existing in other, less structured data sources.

Machine learning models, particularly those capable of processing unstructured data - text from public filings, payment pattern data from transaction records, activity on public registries, even signals from publicly available news and regulatory disclosures - enable significantly broader insight into a business partner's health than traditional models can offer.

This expanded data breadth carries dual value. First, it enables assessment of partners who would otherwise remain beyond the reach of traditional models due to lack of long history. Second, and perhaps more importantly, it enables earlier detection of deterioration among existing partners - AI models capable of tracking subtle changes in payment patterns, tone of public communications, or public registry activity can signal rising risk months before that deterioration would show up in traditional, quarterly-updated financial indicators.

It is worth noting that this early detection capability does not replace traditional analysis, but complements it. Organisations treating AI models as a complete replacement for human judgement, rather than an additional layer of information informing that judgement, risk a new kind of error - excessive reliance on a black box whose internal decision-making is not always transparent or explainable.

This brings us to one of the most important practical challenges in applying AI to credit risk assessment: the demand for explainability. When a traditional statistical model assesses risk as elevated, an analyst can relatively easily trace which specific factors - say, a liquidity ratio falling below a certain threshold - led to that assessment. With more complex machine learning models, particularly those based on deep neural networks, that path from input data to output risk assessment can be considerably less transparent.

This lack of transparency is not merely an academic concern. Regulatory frameworks across multiple jurisdictions increasingly explicitly require organisations to explain why a particular business partner was denied credit or a business relationship, especially when that decision significantly affects that partner's operations. A model unable to offer a comprehensible explanation of its own assessment creates regulatory risk regardless of how statistically accurate it may be.

This has spurred development of an entire subfield known as "explainable AI" (XAI), which attempts to build models whose decisions can be understandably translated into language a human analyst, and eventually a regulator, can follow and verify. Organisations selecting AI tools for credit risk assessment today should treat explainability as an equal criterion alongside pure predictive accuracy, not a secondary consideration addressed after the fact if a regulator raises a question.

One of the most valuable practical shifts AI enables is the transition from periodic to continuous credit risk monitoring. Traditionally, a business partner's credit risk assessment occurs at discrete points - when establishing the relationship, perhaps annually at contract renewal, or reactively when a concrete problem signal appears.

This periodic approach creates natural gaps in oversight - deterioration occurring between two scheduled assessments goes unnoticed until the next scheduled review, which could be months away. AI systems capable of continuously processing available signals - new public filings, changes in transaction patterns, public registry updates - enable a shift toward a model where risk is tracked nearly in real time, with automatic alerts when a defined concern threshold is crossed.

This shift also carries organisational implications extending beyond the technology itself. Continuous monitoring generates a significantly higher volume of signals than periodic review, requiring clearly defined escalation thresholds - not every mild signal deserves immediate human analyst intervention, since that would quickly lead to alert fatigue that ultimately reduces, rather than increases, the overall effectiveness of the monitoring system.

The technical capability of an AI model to generate precise, timely credit risk signals is worth little if the organisation lacks a clear structure for converting those signals into concrete action. This is an area where many organisations, despite significant investment in the AI technology itself, fail to realise the full value of that investment.

The practical question is: who is responsible when an AI system signals elevated risk at a key business partner? Does that person or team have clear authority to act - halting further transactions, requiring additional guarantees, escalating to senior management - or does the signal simply remain logged in the system without a clear owner to act on it?

Organisations most successful at integrating AI into credit risk management are typically those that clearly defined the organisational process before introducing the technology - who receives signals, what criteria they use for prioritisation, what response speed is expected, and to whom escalation occurs when a signal crosses a certain severity threshold. Technology without this organisational framework creates merely an additional data source, not a genuinely improved risk management system.

There is a subtle but real risk accompanying successful AI system implementation for credit risk assessment: gradual erosion of human critical judgement in favour of unquestioning trust in the model. When an AI system demonstrates consistently good accuracy for months or years, there is a natural tendency for human analysts to gradually stop questioning its recommendations, treating them as nearly infallible.

This pattern becomes particularly dangerous precisely at moments when critical human judgement is most needed - during unusual, structurally novel economic conditions for which the model may not have been trained. Machine learning models, however sophisticated, ultimately learn from historical data, and their ability to predict behaviour in genuinely novel, unprecedented situations remains uncertain.

Organisations wanting to avoid this risk must actively cultivate a culture treating AI recommendations as valuable but not infallible input, maintaining regular human verification practice - particularly for high-value or high-risk decisions, where the cost of misjudgement significantly outweighs the cost of additional time spent on human verification.

The regulatory framework governing AI use in financial decision-making is undergoing active evolution worldwide. The European Union, through its comprehensive approach to artificial intelligence regulation, classifies systems affecting credit access as high-risk applications, subject to stricter transparency, documentation, and human oversight requirements than low-risk AI applications.

Organisations implementing AI systems for credit risk assessment today should treat these regulatory requirements not as an obstacle to be circumvented through minimal compliance, but as a useful framework that, if genuinely applied, naturally leads toward more robust, reliable systems - documentation enabling a regulator to understand the model is the same documentation enabling an internal team to detect and correct problems before they become serious.

For organisations considering or already beginning AI integration into credit and counterparty risk assessment, several practical principles emerge from the experience of organisations that have already traversed this path. First, model explainability deserves equal attention to its predictive accuracy, not a secondary role addressed only when a regulator raises a question.

Second, continuous monitoring requires an equally carefully designed escalation process as the model itself - technical capability to generate signals without a clear organisational path for acting on those signals creates false security without genuine risk management improvement.

Third, human critical judgement must remain an active, not passive, part of the process, particularly for high-value decisions - this requires a deliberate organisational culture actively encouraging questioning of AI recommendations, not merely a formal policy permitting this on paper while practice gradually slides toward unquestioning acceptance.

GNK ASG d.o.o. monitors developments in this area as part of a broader approach to financial risk management within the GNK DINAMO Ltd. Group, in the conviction that AI represents a valuable but insufficient tool for credit risk assessment - genuine value stems from thoughtful integration of technology with robust organisational processes and sustained human critical judgement.

While precisely quantifying the advantage of AI-assisted risk assessment remains challenging due to contextual differences between organisations, available data suggests significant, measurable benefits for organisations that have thoughtfully implemented this transition. Early detection of partner creditworthiness deterioration - months before traditional, quarterly-updated indicators would signal it - directly translates into reduced exposure to bad receivables.

Equally important, expanded data breadth enables organisations to enter business relationships with partners who would otherwise remain beyond reach due to lack of long credit history - opening business opportunities a more conservative, exclusively traditional approach simply could not recognise as acceptable.

This dual benefit - reduced risk alongside expanded business opportunities - explains why an increasing number of organisations, despite genuine explainability and regulatory compliance challenges, continue investing in this technology as a long-term component of their risk management infrastructure.

The discussion so far has focused predominantly on assessing direct business partners - clients, suppliers with whom the organisation directly contracts. But AI also opens the possibility of extending risk assessment deeper into the supply chain, toward suppliers' suppliers, whose health directly affects the reliability of the direct partner, but traditionally remains entirely outside the organisation's field of view.

This capability - mapping and assessing risk across multiple supply chain tiers - becomes increasingly relevant as global supply networks grow in complexity, and geopolitical and climate risks increasingly cause disruptions originating deep within the chain, beyond the direct field of view of the organisation that ultimately feels the consequences.

Organisations investing in this expanded visibility, though requiring more significant initial investment in data collection and processing, position themselves for better prediction and mitigation of disruptions that would otherwise appear as complete surprises, despite their early signals having existed deep within the supply chain months before materialising as a visible problem for the organisation itself.

The final lesson emerging from the experience of organisations most successful at integrating AI into credit and counterparty risk management is that technology functions as an amplifier of existing organisational discipline, not a substitute for it. Organisations with already robust risk management processes see AI as a tool making those processes faster, broader, and more precise. Organisations lacking fundamental risk management discipline risk AI simply automating and accelerating existing weaknesses rather than correcting them.

This sets a clear priority sequence for organisations considering this transition: before investing in sophisticated AI technology, it is worth ensuring that fundamental organisational processes - clear decision ownership, defined escalation thresholds, a culture nurturing critical judgement rather than unquestioning acceptance - already exist or are being actively built in parallel with the technological investment.

#ArtificialIntelligence #CreditRisk #GNKASG #NerminSefic #GNKDINAMOLtd #SeficNermin


Cjelovit tekst i izvor: https://gnk-asg.hr/en/analyses/artificial-intelligence-credit-counterparty-risk/

Autor i urednička odgovornost: Nermin Sefić. Izdavač: GNK ASG d.o.o..

#GNKASG #GNKDINAMOLtd #NerminSefic #BusinessIntelligence #artificialintelligence #creditrisk #counterpartyrisk #NerminSefić #explainableAI

Halucinacije AI modela

Autor: Nermin Sefić Halucinacije velikih jezicnih modela dolaze u obliku identicnom tocnom odgovoru - visesloevita strategija ublazavanja ...