From problem to evidence: how to design AI pilots in tax administrations

The pilots allow to determine if AI improves a tax task compared to the current process, with acceptable benefits, costs and risks.

 

Key ideas:

* The value of AI in tax administrations depends on the problem it solves and whether it offers better results than the current process, with acceptable costs and risks.
* The pilots allow to check it in a controlled environment, comparing performance, costs and risks and verifying if the safeguards work in practice.
* A successful pilot provides evidence to decide if it is appropriate to discard, adjust, expand or prepare for scaling a solution.

Tax administrations have been using advanced analytics, machine learning and artificial intelligence for years to detect risks, select cases, answer queries and automate tasks.

In 2023, 69% of the administrations analyzed by the OECD had implemented and used AI, including machine learning, while another 24.1% were in the process of implementation. In Latin America and the Caribbean, data from the Inter-American Center of Tax Administrations (IATTC), show an important difference between technology: in 2022, among 34 tax administrations, 70.6% used or were implementing data science and analytics tools, while AI was in use or in implementation in 23.5%. These indicators cover different AI applications and do not specifically reflect the use of generative artificial intelligence.

The emergence of generative AI marks a new step in this technological evolution, expanding opportunities to analyze documents, generate drafts, answer queries and facilitate access to institutional knowledge. But it also introduces specific risks, such as made-up answers, disclosure of confidential information and dependence on external suppliers.

The pilot allows to evaluate these opportunities and risks before integrating the solution into the operation. In a limited environment, it compares its performance with the current process, identifies faults, estimates costs and tests safeguards.

Your central question is: does this task improve compared to the current process, with acceptable costs and risks? To answer it, it is necessary to define six elements.

 

1. Choosing a problem, not a tool

The starting point should be an observable bottleneck: delayed files, repetitive consultations or an imprecise selection of cases. “Testing a chatbot” describes a solution, not the problem. The team must define a task, identify the user and measure the baseline: time, cost, quality and errors of the current process.

Not every problem requires a language model. Structured data can be solved better with rules or predictive models; documents, with specialized classifiers; and the summary or the consultation of regulations, with language models connected to institutional sources. The rule is to choose the least complex alternative that allows to achieve the defined results and thresholds.

For [user], we will improve [result] by [assisted task], comparing it with [current process] during [period], without delegating [reserved decision to a person].

 

2.Define what the AI will do and who will respond

The AI can retrieve information, prepare a draft, formulate a recommendation or execute an action. The more you influence an inspection, refund or sanction, the greater the demands for explanation, traceability and human review must be. Supervision is only effective if the person has the information, time and authority to question or reject the result.

These requirements must also be incorporated into the architecture: authorized sources, permissions, limits of action and mechanisms to correct or interrupt its operation. Therefore, once the task and responsibility have been delimited, it is necessary to examine the entire system, that is, model, data, tools, integrations and controls, and not only the model used.

An institutional solution combines the model with knowledge sources, a platform that manages access and records, the application and its controls. An agent adds autonomy by choosing steps and invoking authorized tools; therefore, the risk also depends on the data it queries, the actions it executes and the permissions it receives.

A small language model (SLM) may be sufficient for bounded tasks; others may require a large model (LLM). In generative applications based on institutional knowledge, augmented generation by retrieval (RAG) can connect the model with approved sources without training it from scratch.   The alternatives — commercial, open or deployed on own infrastructure – should be compared with the same cases and evaluated according to quality, invented responses, latency, total cost, privacy and auditability.

The Generative AI Profile of the NIST it provides guidance for evaluating the performance and managing risks of generative AI. The goal is to choose the simplest architecture that meets the thresholds of the case. However, no architecture is viable if the data that feeds it is inadequate or exposed.

 

3. Protect the data before connecting it

Before using data, its origin, quality, representativeness, permissions and restrictions must be verified. When possible, the first tests should use synthetic or anonymized data, institutional accounts and independent cases to evaluate the system.

If a provider is involved, it should be specified what information is sent, where and for how long it is processed and stored, who can access it and how it will be deleted. Not using the data to train the model does not mean that it will not be retained. The NIST Generative AI Profile recommends protecting data and managing risks arising from third parties. The legal basis, minimization, access controls and applicable international transfers should also be verified with real data.

Once the data is protected, it is necessary to demonstrate whether the solution provides value.

 

4. Agree on what success means

Technical precision is not enough. Before starting, the pilot must establish a baseline, indicators and thresholds for continuing, adjusting or stopping the four-dimensional solution:

  • Operation: time, cost and rework.
  • Quality:accuracy and correct answers.
  • Risk and equity:false positives, false negatives and incidents.
  • Adoption:effective use, corrections and user satisfaction.

A the Brazilian case  shows why. The first tests of an AI project to support tax litigation management used 2,000 manually tagged files and reached a sensitivity and specificity of more than 80%. This is a promising technical result, but it does not in itself demonstrate savings, fairness or a better quality of the final decision.

That is why, in addition to measuring results, the pilot must reveal when it works and how it can fail.

 

5. Also test how it fails

The initial environment should limit users, data, duration and consequences. Whenever possible, it is advisable to start in “shadow mode”: the system generates results, but does not intervene in real cases while comparing with the current process.

The tests should include both normal and adverse situations. In predictive models: incomplete data, changes in the data and errors between groups. In generative AI: made-up statements, malicious instructions, disclosure of information, inconsistent answers and incorrect sources.

The team must integrate the user, data, security, privacy, legal affairs and internal control areas. The the NIST framework for AI it organizes this work into four functions: governing, mapping, measuring and managing. The tests should document results, incidents and observed limits. That evidence will serve to determine the next step.

 

6. Close with a decision, not a demonstration

The pilot must conclude with evidence that brings together the baseline, results, costs and incidents, user experience, limitations and residual risks. With it, the institution must decide whether to stop the initiative, redesign the solution, expand the pilot in a limited way or prepare its escalation.

Preparing for scaling is not an automatic consequence of a good technical result. A pilot demonstrates limited feasibility; bringing the solution to the operation requires integrating processes, architecture, talent, purchasing, budget, governance and permanent monitoring, capabilities highlighted by the OECD when analyzing the passage of AI pilots to its implementation.

Therefore, these conditions must be agreed upon from the very beginning. A good pilot does not end up proving that the technology works, but determining whether it generates enough value to justify the next step.

 

The form that must be approved

Before starting, the business, technology, data, security and compliance teams must agree on the questions that the pilot will answer. The sponsor must approve a short form summarizing the objectives, assumptions and evaluation criteria of the initiative:

 

 

A pilot must answer whether AI is the best alternative 

AI offers a great opportunity to improve the efficiency and effectiveness of tax administrations. However, its potential does not eliminate a fundamental reality: its adoption does not necessarily imply better results.

The challenge is not to multiply demonstrations, but to turn institutional problems into comparable and auditable tests on when to use AI, for what, with what limits and under what conditions. Only then the second conversation begins: how to move from pilot to an institutional capacity without losing control, legitimacy and public purpose.

This article was reproduced with the permission of the author Monica Calijuri, originally published in the the IDB’s blog.

Disclaimer. Readers are informed that the views, thoughts, and opinions expressed in the text belong solely to the author, and not necessarily to the author's employer, organization, committee or other group the author might be associated with, nor to the Executive Secretariat of CIAT. The author is also responsible for the precision and accuracy of data and sources.

Leave a Reply

Your email address will not be published.

CIAT Subscriptions

Browse through the site without restrictions. Consult and download the contents.

Subscribe to our electronic newsletters:

  • Blog
  • Academic offer (Only in spanish)
  • Newsletter
  • Publications
  • News alert

Activate subscription

CIAT Members

Representatives, Correspondent and Authorized staff (TA)