How artificial intelligence supported a deeper analysis of FAO’s work for small-scale producers
©FAO/Andrew Esiebo
In a recent evaluation of FAO’s support to small-scale producers’ equitable access to resources (BP4), artificial intelligence (AI) assisted tools helped the evaluation team process and analyse large volumes of evidence while maintaining methodological rigour and human oversight.
The tools supported with interview transcription, qualitative coding, comparative analysis, pattern identification and editorial review. Their most structured use focused on portfolio analysis, where AI helped classify reporting narratives covering work undertaken across FAO headquarters and Decentralized Offices.
The stocktaking exercise covered 632 entries related to the Programme Priority Area better production 4. The entries described planned interventions and reported achievements across a wide range of countries, technical areas and organizational units.
Anyone who has worked with corporate reporting data will recognize the challenge. The source material varied considerably. Some narratives were short and specific. Others were longer and covered several activities or thematic areas. More than 140 had originally been submitted in French or Spanish.
Before using AI, the evaluation team cleaned and standardized the data and translated the non-English entries. It then developed a codebook covering four dimensions: thematic areas of work; inclusion; acceleration mechanisms; and FAO’s core functions. An AI model was instructed to apply these categories to each narrative. However, initial testing showed that the first version of the methodology was too restrictive.
The original rules were designed to retain only entries that showed clear mechanisms or outcome-level change. This excluded information that was relevant to the purpose of the exercise, which was to map the scope and characteristics of FAO’s work rather than assess the effectiveness of each individual intervention.
The initial approach also allowed only one label per narrative. Testing showed that many entries covered several themes or functions.
The team therefore revised both elements. It removed the exclusion criteria and introduced multi-label classification, allowing several categories to be assigned to the same entry.
These methodological decisions shaped the analysis more than the choice of model.
The coding framework was developed by the evaluation analyst and the lead evaluator, with inputs from programme leads, subject-matter experts and staff of FAO’s Office of Strategic Planning and Budget. The model applied the framework, but evaluators defined the categories, clarified areas of overlap and decided how the results should be interpreted.
Making portfolio-wide comparison possible
The resulting dataset allowed the evaluation team to compare entries across the portfolio and identify patterns in the four coded dimensions. These patterns were then examined alongside evidence from document review, interviews and case studies.
Coding the 632 narratives was therefore only one part of the exercise. Its main benefit was to make portfolio-wide comparison more manageable and to provide a structured basis for further analysis.
The workflow was semi-automated. A script prepared the data, submitted each narrative to the model and compiled the results. Validation, adjudication and interpretation remained under human control.
Testing the classifications
The team tested the reliability of the coding process before and after the full portfolio analysis.
Before the formal analysis, an independent parallel coding exercise was conducted to test the applicability of the codebook and the AI-powered workflow. A small sample was coded independently by a human reviewer and the primary AI model. The team then examined areas of agreement and disagreement and proceeded with the formal analysis only after the results were considered satisfactory.
Following the full analysis, a second AI model and a human reviewer jointly assessed a stratified sample of 60 narratives, representing approximately 10% of the total sample. The sample deliberately included difficult cases, such as unusually long or short entries, narratives containing several thematic signals and texts with vague or ambiguous wording.
After this review, the primary model required revision in only one of the 60 cases. The two AI models had produced partially different combinations of labels in 15 cases, but many of these differences involved an additional or omitted secondary label rather than a fundamentally different interpretation of the narrative.
The results gave the team confidence that the classifications were sufficiently consistent for the purpose of descriptive portfolio mapping. They also reinforced the importance of targeted human review where source text is ambiguous, several themes overlap or category boundaries are difficult to distinguish.
AI support beyond portfolio coding
Portfolio classification was the most fully documented AI-assisted component of the evaluation, but it was not the only one.
AI tools also assisted with interview transcription, qualitative coding, comparative analysis, pattern identification and editorial support to enhance clarity and consistency in drafting.
These uses carried different levels of analytical significance. In every case, evaluators remained responsible for designing workflow and protocols, checking outputs, assessing the quality and relevance of the evidence and deciding what could support a finding, conclusion or recommendation.
The tools supported specific tasks within the evidence-processing workflow and helped evaluators manage the volume of material.
The exercise also followed institutional requirements for responsible AI use. All internal data were processed using AI services subscribed to by FAO, including Microsoft Copilot, and OpenAI GPT models deployed through the Azure environment. They were not used for external model training or retained beyond the evaluation exercise. The approach followed relevant FAO and United Nations Evaluation Group guidance on the responsible use of AI.
Three lessons for the evaluation community
The experience highlights three practical lessons.
- Start with the analytical question, not the technology. In this case, the purpose was descriptive: to map the scope and characteristics of FAO’s work, not to assess the effectiveness of each intervention.
- Invest in the analytical framework. Clear definitions, examples and decision rules were essential. Pilot testing revealed where the original approach was too restrictive and where multi-label classification was needed.
- Build validation into the workflow from the start. A second model can help identify uncertain cases, but it cannot replace human review and adjudication.
Used this way, AI can help evaluation teams work more systematically with large volumes of evidence. Its value depends not only on the technology, but on the quality of the framework, the validation process and the evaluators responsible for interpreting the results.
Read more about the methodology for the AI-assisted descriptive analysis