Trust and safety at scale: Can automation and AI deliver without human oversight? 

As digital platforms expand, automation and AI are revolutionizing content moderation and trust and safety operations. But with challenges around bias, transparency, and accountability, can these technologies truly deliver effective results without the vital input of human moderators?

Overhead view of a man sitting at a desk with a laptop and large monitor in front of him

Published ·August 24, 2026

Reading time·8 min

Automation and artificial intelligence (AI have had a transformative effect on trust and safety — especially when it comes to content moderation. Because, as the size, complexity, and popularity of digital platforms continues to grow, these tools have enabled trust and safety teams to keep pace and scale operations, maintaining a secure and reliable environment for users when assessing digital services or engaging in transactions.  

We have arrived at a point in time when it is simply not possible to execute a robust, multi-region (let alone global) trust and safety policy underpinned solely by manual moderation. And it’s not simply because of the sheer volumes of content requiring review, it’s also because the risks to reputation and even brand survival when moderation efforts can’t keep pace are significantly greater.  

AI and automation tools currently play a very specific role within trust and safety delivery. However, as the scalability and proactivity of trust and safety operations become more reliant on these technologies, and as the hype around the potential applications and capabilities of generative and agentic AI continue to grow, so does the belief in some quarters that we will soon be able to fully automate all aspects of enacting a trust and safety policy.  

But to understand how continued technological advances could benefit current trust and safety efforts and to assess the feasibility of the claim that trust and safety could soon become an end-to-end automated process, without human intervention, we need to first understand the realities of trust and safety delivery.  

The reality of trust and safety automation today 

Within trust and safety, “AI and automation” is an umbrella term that covers any number of tools and solutions that work together to execute certain clearly defined tasks. Those tasks and tools can vary from organization to organization depending on its size, sector, userbase and even its geographies. However, in general, these technologies are applied and excel in three areas:  

Identify: 
By checking for matches with known abuse, following set rules, or using past examples to make predictions, AI and automation actively detect harmful content or behavior in the moment.  

Act
When content or behavior that clearly goes against a trust and safety policy is identified, automated systems can also immediately take the appropriate action as defined by the policy. So, this could mean reducing the content’s visibility, removing it from the platform, or suspending or deleting a user’s account.  

Ask
When, due to the limits of their existing data or training parameters, the tools are unable to decide if content or behavior violates the trust and safety policy, they also route content to human moderators for review.  

In other words, AI and automation excel at handling clearly defined, repetitive moderation tasks on a large scale, allowing for efficient processing of high content volumes. This not only streamlines operations but also helps reduce the amount of potentially disturbing content that needs to be escalated for manual human review. 

Current automation limitations 

But, as well as being focused on specific tasks, these technologies are always applied as elements of a multi-layered approach to trust and safety, the idea being that no layer (i.e., process, solution or tool) is either foolproof or able to focus on every vulnerability. Therefore, if one layer fails, other layers still protect the system — hence the need for human review of nuanced or novel cases, the ability for other users to flag or report content, plus the fact that the best trust and safety policies always have a clear appeal process for instances where a user or community of users wish to challenge a decision.  

Trust and safety best practices and the technologies used to support them pre-date the arrival of generative AI. This is a discipline that has been developing organically since the turn of the millennium. As such, the capabilities of current tools and solutions have been built up slowly over that time and often required significant financial or human resources. For instance, an AI solution that, through machine learning (ML), is able to rapidly classify different types of content as being acceptable or unacceptable must first be trained on data specific to the company, its policy and the user base. And this means the manual curation of thousands of correctly labelled examples of content that is or is not acceptable.  

Once initial training is completed, the solution still needs to be refined constantly so that it can meet an agreed level of performance and so that it can identify and classify new types of content.  

For the same reasons when automated solutions are used to support trust and safety activities, the goal is to find a balance because no system is 100% effective or flawless. And typically, that balance is between over-performing and underperforming. If the solution flags too many pieces of content or types of behavior that don’t go against policy, it is over-performing or producing too many false positives. And if it constantly fails to identify instances where policy is being violated, it is underperforming or producing too many false negatives.  

Why there’s excitement about GenAI 

If we consider this historical approach, it’s little wonder why there is excitement about the capabilities of generative AI in respect to content moderation. Whereas ML solutions can only guess if something should be flagged or removed based on historical data or evidence, the way GenAI systems are created means they can process images and language like a human, even before they are calibrated for specific tasks.  

As such, creating tools that can spot and classify instances of unacceptable content or behavior would be much faster, cheaper, more efficient and potentially more accurate, further reducing the need for human moderators to review anything other than edge cases.  

Why there are still limitations 

All AI models have biases and limitations, and generative AI is no exception. And these biases can manifest in a number of ways: 

1. Training data 

The quality and quantity of data used to train an AI model will have a direct influence over its potential for delivering biased results. For instance, if it is trained exclusively on English-language sources, outputs could be perceived to reflect a western bias. Or, if a model was trained on a much wider geographical sample of data but each piece of training text pre-dated 2001, outputs could appear to reinforce or prefer outdated societal or political views.  

2. Algorithm design 

The algorithms themselves can introduce bias if they are not designed with fairness in mind. For example, choices made regarding how to weigh different types of input data can create an imbalance that favors certain outcomes over others, irrespective of the training data.  

3. Feedback loops 

Generative AI relies on user feedback to judge, confirm and improve its outputs. If the feedback mechanisms aren’t properly weighted or tested for fairness or objectivity, then they can create and reinforce inequalities in the model during and after development so that its outputs are biased.  

Transparency issues 

Trust and safety demands not only an unbiased but also a transparent approach — it’s a key tenet of the discipline insomuch as teams need to be able to articulate how and why they took a certain action. However, with very few exceptions, GenAI systems are opaque — they cannot explain how they arrived at their answer or solution, and this lack of clarity and transparency can lead to a lack of trust and can compound issues around bias and fairness. 

1. Interpretability 

As many AI systems are essentially black boxes where a question goes in and an answer comes out, but the calculations to arrive at the response are hidden, it makes it even more difficult to identify if a model is biased and in which regard. And if users can’t look under the hood, so to speak, they are going to be less likely to trust it or its outputs.  

2. Accountability 

And of course, if there is no way of understanding the decision-making process, it becomes difficult to pinpoint where the problem lies — in the training data, model design, user input or initial calibration through feedback mechanisms. Without knowing who or what is accountable, it is difficult to develop or establish a governance structure for the system’s use.  

User perception 

These issues can be further compounded by external factors, the biggest of which being how a platform or digital business that’s over reliant on AI is perceived by its users, members or customers. There is growing concern, bordering on distrust, of advanced AI systems at a societal level. Without transparency, users are likely to judge actions taken by a trust and safety function as arbitrary or unfair. Likewise, even generative AI can struggle with nuance or context when analyzing content and, without human oversight, this can lead to too many false positive results, which again can lead to distrust or negative publicity. And, if left unchecked, could drive users to competitors.  

Moving to full automation can also create regulatory risks. As a means of keeping AI in check, many governments around the world are rolling out legislation — such as the EU’s Digital Services Act — which are focused on making the technology’s use transparent and accountable. And if the trust and safety function is fully automated, there is a strong possibility that it will run afoul of regulations.  

A hybrid approach 

Despite all these issues, AI and automation are already making a positive, sustained difference within trust and safety by shielding human reviewers from abusive or disturbing content. This, in turn, enables moderators to apply their own unique abilities — the understanding of nuance, sarcasm, cultural context and empathy to the decision-making process to ensure fair judgement. 

And, as the use of and the capabilities of AI-based solutions expands, so will the need for human oversight for calibrating performance, continued training and adjustment as new risks or types of policy violation are identified.  

Digital platforms and services have become an integral element of daily life across diverse demographics and geographies. As such, maintaining their integrity demands a balanced approach so that engagement and expression are not sacrificed purely for the sake of eliminating risk.  

The diversity of both user bases and the platforms themselves means that there is currently no one-size-fits-all solution to developing and executing the right trust and safety policy. Every approach or framework will be different except in one respect, its flexibility to change over time in line with evolving risks and new threats. And this ability to change and evolve is only possible when automation is paired with human oversight.   

To learn more about trust and safety, read our ebook: “Moving beyond the myths: The reality of trust and safety in modern business.”