As the Caribbean nation of Grenada accelerates its shift from conceptual discussions of digital transformation to tangible on-the-ground implementation, artificial intelligence has emerged as a central pillar of the government’s modernization agenda. The country’s Information and Communication Technology (ICT) Division has already elevated AI-powered public administration tools to its list of top strategic priorities, aligning with broader trends across the Caribbean documented at the recent CANTO (Caribbean Association of National Telecommunication Organisations) conference, which painted a picture of a rapidly evolving regional digital economy increasingly centered on AI innovation. Later this month, Grenada will further cement its role in regional digital development as a co-host of the Caribbean Internet Governance Forum, an event where digital transformation strategies and coordinated regional tech policy will take center stage in discussions.
This growing momentum for AI adoption opens up significant new opportunities to improve public service delivery, streamline government operations, and boost Grenada’s digital competitiveness. However, it also carries a well-documented risk common to new technology pilot projects: untested initiatives often become permanently embedded in bureaucratic systems not because they deliver proven public value, but simply because they have been implemented and institutionalized over time.
To address this critical governance gap, behavioral scientist Gleb Tsipursky, PhD, is calling for all public sector AI pilots in Grenada to adopt a mandatory “sunset test” framework before any tool is launched. Under this approach, government agencies are required to set a fixed formal review date before launch, explicitly document the specific public problem the AI tool is intended to solve, establish a baseline measurement of current performance, and outline clear pre-defined criteria that will determine whether the pilot is expanded, modified, or terminated after the review period.
The specific metrics for a sunset test can be tailored to the function of each AI tool, Tsipursky explains. For example, a citizen-facing public service chatbot would be evaluated based on metrics such as the share of user inquiries that receive a complete, accurate answer on the first interaction (eliminating the need for follow-up contact with human staff), the frequency of incorrect responses that require staff correction, and the types of questions that still demand human intervention. For internal AI drafting tools used by government employees, success would not be measured by the speed of the first draft alone, but by the total net time saved after accounting for required staff review and revisions. For AI systems designed to prioritize public case processing or application reviews, key metrics would include how often human staff override the AI’s ranking recommendations and the core reasons for those overrides.
Tsipursky emphasizes that sunset testing is non-negotiable because AI systems tend to accumulate hidden, long-term costs that make them difficult to remove even when they fail to deliver on promised value. Over time, employees develop workarounds to compensate for the tool’s flaws, managers restructure entire work processes around the system, and external AI vendors become deeply embedded in government operations. Within just six months of launch, an experimental AI pilot can become nearly impossible to discontinue, even when there is little hard evidence that it actually improves public outcomes.
This framework does not require Grenada to reject imperfect or experimental AI pilots, Tsipursky notes – intentional experimentation is a necessary part of institutional learning and digital innovation. Instead, it establishes a clear rule that every pilot must prove its public value to earn a permanent place in government infrastructure.
Conducting a useful sunset test review does not require lengthy, bureaucratic reports: Tsipursky notes that a rigorous review can be completed in a single page, focused on six core questions: What measurable public outcomes improved after implementing the tool? What new unplanned work did the tool create for government staff? What persistent errors continued to occur even with the AI system in place? Where did human judgment remain irreplaceable for delivering high-quality outcomes? Did AI implementation reduce delays for citizens, or did it create additional confusion? And finally, what tangible impact would occur if the tool was removed tomorrow?
That final question is particularly critical, Tsipursky argues. If an AI pilot has truly become an indispensable part of public service delivery, government leaders deserve a clear explanation of why it creates value. If no team can identify a measurable loss of public value from turning the pilot off, the initiative is almost certainly wasting limited government attention and funding that could be allocated to higher-impact digital projects.
As a small nation advancing its digital agenda at a rapid pace, Grenada has a unique opportunity to balance speed with disciplined governance, Tsipursky concludes. A mandatory sunset test framework makes experimentation politically and administratively defensible, because every new AI tool is introduced with clear evidence standards and a built-in exit strategy for underperforming initiatives. This approach is the most effective way for Grenada to move fast on digital transformation without allowing unproven yesterdays experiments to become unchallenged, wasteful dependencies tomorrow.
