When One Character Disrupts a Harvest — Exchanging Experiences on Metadata Encoding During the FAO AGRIS Brown Bag Session

08/07/2026
 When One Character Disrupts a Harvest — Exchanging Experiences on Metadata Encoding During the FAO AGRIS Brown Bag Session

Pexels

The FAO AGRIS Brown Bag Session brings together repository managers and data providers from organisations around the world to exchange experiences on metadata management and interoperability — addressing challenges that many institutions face individually but rarely have the space to discuss collectively. 

One of the issues raised in a recent session was the handling of different alphabets and special characters in titles, abstracts and metadata fields, a challenge that reflects the richly multilingual nature of the FAO AGRIS network itself. Among the approaches discussed, UTF-8 emerged as a key reference point — a widely adopted encoding standard that quietly enables repositories around the world to communicate using the same technical language.

Repository managers sometimes face difficulties with character encoding, often only noticing the issue when something goes wrong — for example, when an author's surname suddenly appears as "García" instead of "García." A title written in Portuguese, Arabic or Chinese becomes unreadable after harvesting. Although the research itself remains intact, corrupted metadata can reduce discoverability, affect search results and compromise the user experience.

UTF-8 is the world's most widely adopted character encoding standard. It allows computers to correctly represent letters, symbols and scripts from virtually every language. Today, it underpins websites, digital repositories, XML documents and metadata exchange protocols such as OAI-PMH, making it an essential component of infrastructures like FAO AGRIS, where research is harvested from hundreds of institutions across multiple languages.

As agricultural knowledge becomes increasingly global, preserving the integrity of multilingual metadata becomes equally important. Repository records often include author names, institutional affiliations, titles, abstracts and keywords containing accented characters or non-Latin scripts. When encoding is inconsistent, these characters may be displayed incorrectly, lost during processing, or may even prevent the correct indexing and retrieval of information. Furthermore, the lack of uniform encoding makes it difficult to exchange data between systems, affects interoperability and can lead to errors in the display, search and processing of records.

Implementing UTF-8 consistently across repository systems offers clear benefits. It preserves the accuracy of multilingual metadata, improves interoperability with FAO AGRIS and other aggregators, and ensures that scientific outputs remain discoverable regardless of the language in which they were produced. These small technical decisions ultimately strengthen the visibility and accessibility of agricultural research worldwide.

Like metadata itself, UTF-8 is easy to overlook when everything works as expected. Yet it is one of the invisible standards that allows knowledge to travel accurately across repositories, countries and languages, ensuring that valuable scientific information reaches the people who need it most.

Sign up

To receive newsletters about FAO AGRIS and FAO Knowledge Management activities

NEWSLETTER

RENEW YOUR SEAL OF RECOGNITION FOR 2027

Institutions that submit data between July 2026 and June 2027 are invited to request the renewal of their FAO AGRIS Seal of Recognition for the year 2027.

 

Seal of recognition for active Agris Data Providers 2024

REQUEST SEAL 2027

MEET THE DATA PROVIDERS

Interested in becoming a FAO AGRIS data provider? Click here