Language Representation in Alma CZ Bibliographic Records Using ISO 639-2 and ISO 639-3 Codes
Following the February 2026 Alma release, support for ISO 639-3 language codes was introduced as part of a combined controlled vocabulary for languages.
As a result, Alma Community Zone (CZ) bibliographic records may include multiple language codes representing the same language, using both:
- ISO 639-2 (MARC standard)
- ISO 639-3 (expanded language set)
This article explains the expected behavior and its impact.
Background
The February 2026 enhancement introduced:
- A merged language vocabulary combining:
- ISO 639-2
- ISO 639-3
- Usage in:
- Language fields within bibliographic records
- Metadata Editor
- Discovery indexing and facets
This enables more granular and comprehensive language representation in Alma.
Behavior in CZ Bibliographic Records
In Alma CZ bibliographic records, it is possible to encounter:
- Multiple language entries representing the same language
- Different ISO standards used in parallel
Example
A CZ bibliographic record may include:
Language: fre
Language: fra
or:
Language: ara
Language: arb
Explanation
- CZ records aggregate metadata from multiple sources and enrichment processes
- These sources may use:
- ISO 639-2 codes (e.g.,
fre,ara) - ISO 639-3 codes (e.g.,
fra,arb)
- ISO 639-2 codes (e.g.,
Alma preserves all values present in the CZ bibliographic record.
Expected Behavior
This behavior is expected and by design.
- Language fields in CZ bibliographic records may contain multiple values
- Alma does not deduplicate or merge language codes
- All language values are:
- indexed
- searchable
- exposed in discovery
Even if the values represent the same language at different levels of granularity
Impact on Discovery
In Primo VE:
- All language values from CZ bibliographic records are indexed
- Language facets may display:
- multiple entries for the same language
- both general and specific forms
Example
| Code | Display |
|---|---|
| ara | Arabic |
| arb | Standard Arabic |
Both may appear simultaneously in the facet.
Why Deduplication Is Not Performed
Alma does not merge or deduplicate ISO language codes in CZ bibliographic records because:
- CZ is designed to preserve source metadata integrity
- ISO 639-2 and ISO 639-3 represent different levels of linguistic granularity
- Automatic merging could:
- remove valid distinctions
- introduce inconsistencies
- impact search and indexing behavior
Recommendations
Institutions working with CZ bibliographic records may choose to:
Accept the default behavior
- Retain both ISO standards
- Benefit from richer language representation
Apply local normalization
- Standardize language codes (e.g., ISO 639-2 only)
- Map or remove duplicate representations
- Implement via:
- normalization rules
- local workflows
Adjust discovery behavior
- Use Primo VE normalization rules to:
- merge or collapse language facet values
- reduce duplication in display
Important Notes
- CZ bibliographic records may include mixed language standards
- No automatic alignment is performed between:
- different ISO standards
- different language fields
- This behavior applies specifically to CZ-managed metadata
Conclusion
In Alma Community Zone (CZ) bibliographic records, the presence of both:
- ISO 639-2
- ISO 639-3
for the same language is:
✔️ Expected
✔️ A result of expanded language support introduced in February 2026
✔️ Not automatically normalized or deduplicated
Institutions requiring consistency should implement local normalization or discovery adjustments.
Internal Note
This article addresses a documentation gap regarding:
Language representation using multiple ISO standards in Alma CZ bibliographic records following the February 2026 enhancement

