Using Normalized Subject Headings from CDI
Introduction
CDI subject terms can come from various sources and providers in various styles and formats. Some CDI sources use a subject vocabulary, while in other cases, the CDI subjects are author-provided keywords or subjects that come from various sources (such as from ‘’aggregator’’ databases).
To provide more consistent, deduplicated, and normalized subjects across the CDI index, the existing Subject field has been divided into the following fields:
-
A (Normalized) Subject field includes only subjects matching the CDI subject controlled vocabulary. They are deduplicated and normalized.
-
A Keyword field that includes all subjects that could not be mapped against the controlled vocabulary. These subjects will be considered keywords.
The CDI subject controlled vocabulary uses the preferred terms and variant terms from the following sources:
-
Library of Congress Subject Headings (LCSH) – information at https://id.loc.gov/authorities/subjects.html
-
MeSH – information at https://www.ncbi.nlm.nih.gov/mesh/
-
A small subset of ProQuest thesaurus terms
For more information about the CDI subject controlled vocabulary, see Rules for Adding Subjects to the Normalized Subject Field.
When enabled in Primo/VE, this functionality provides the following:
-
Only the normalized subjects appear for the Subject/Topic facet and the Subject display field.
-
When performing a search in CDI scopes, both the Subject and Keyword fields are searched to ensure that nothing is lost (since not all subjects can be mapped to the controlled vocabulary).
The Subject and Keyword fields are configurable for display in the brief results and full display.
At this time, we will not add subjects/keywords to a record unless it already contains them from one of its original data sources. The source for populating the normalized Subject field remains the content provider’s data.
Enabling CDI Normalized Subject Headings
For information on how to configure this functionality, see the associated page for your environment:
Using the Subject and Keyword Fields
Search and Ranking
Enabling normalized subjects can improve both search accuracy and ranking.
Subject field searching:
-
Subject field search queries are matched against normalized subjects rather than the original subject terms, leading to more accurate and relevant results.
Default (any field) searching:
-
Default (any field) search queries are also matched against normalized subjects leading to more accurate and relevant results.
-
Keywords, derived from original subjects that do not map to normalized subjects, are still included in the search, but they carry less weight in the relevance ranking.
Most differences will appear in how search results are ranked, but in some cases the total number of results may also change. For example, a record may have the subject “Cow’s Milk Allergy”, which is mapped to the normalized subject “Milk Hypersensitivity.” When the normalized subject feature is enabled, a search for “cow” would no longer return that record (unless the record contains "cow" in another field).
Display
When switching to the new field, the Subject field is automatically populated from the normalized subjects. You can configure the new Keyword field for display only.
The Normalized Subject Field
The normalized Subject field uses the CDI controlled subject vocabulary, which consists of subjects and their alternate terms. The subjects are compiled from a combination of the following sources:
-
LCSH (Library of Congress Subject Headings) – The primary source for subjects and their alternate terms.
-
MeSH (Medical Subject Headings) – The mapping file compiled by Northwestern University Libraries to avoid duplicates with LCSH. Mapping information can be found at: https://galter.northwestern.edu/about-us/northwestern-university-libraries-lcsh-mesh-mapping-project
-
ProQuest Thesauri for a select set of additional terms that are very closely matched to LCSH.
When closely related or near-duplicate subjects are identified, LCSH terms are given priority. The deduplication process relies primarily on Northwestern University Libraries’ mapping project, with some further refinements made by Ex Libris.
Examples:
- LCSH includes “Adulthood” (sh85001056), while MeSH lists “Adult” (D00328)
- LCSH uses “Human beings” (sh85080292), whereas MeSH uses “Humans” (D006801).
In such cases, the LCSH terms — “Adulthood” and “Human beings” — are retained, and the corresponding MeSH terms — “Adult” and “Humans” — are excluded.
Rules for Adding Subjects to the Normalized Subject Field
General
The subjects in the source record are normalized by converting uppercase letters to lowercase, removing extra spaces and punctuation, and then mapping the subjects against the CDI subject controlled vocabulary If a subject can be mapped, it will be added to the normalized Subject field.
Some exceptions apply to all vocabularies used. For example, we are not adding very general terms such as Style, Intention, or content types (such as Electronic books). Variants are added as alternate terms for the subject unless they are programmatically determined to be mappings to narrower subjects, mappings to broader subjects, or ambiguous subjects.
LCSH
LCSH is the primary source for the CDI controlled vocabulary subjects and their alternate terms. For example, the LCSH subject Driving of horse-driven vehicles (sh85039614) has the following variants (UF) as shown in https://www.loc.gov/aba/publications/FreeLCSH/D.pdf:
-
Driving
-
Driving, Horse-drawn vehicle
-
Horse-drawn vehicle driving
Because the Driving variant maps from a broader to a narrower subject, it is not used as an alternative term for the subject. The two other variants are used as alternate terms for this subject. This means that if a CDI record contains the subject horse-drawn vehicle driving, it is not replaced with the Driving of horse-drawn vehicles subject.
For LCSH complex subjects, CDI adds all components to the new normalized Subject field that exist as LCSH subjects. LCSH Complex subjects contain multiple components typically connected with a double dash, for example—"Japanese American -- Alcohol Use" (sh2008009026). In this example, both "Japanese Americans" (sh85069603) and "Alcohol Use" (sh99002331) are LCSH subjects, so they are both included in the new normalized Subject field.
This is different from the following example: "Japanese Americans--Forced removal and internment, 1942-1945" (sh85069606). Only its first component "Japanese Americans" (sh85069603) exists as a LCSH subject (and is included in the new normalized Subject field), but because the second component, "Forced removal and internment, 1942-1945," does not exist as an LCSH subject, it is not included in the new normalized Subject field, but instead, it is included with the Keyword field.
MeSH and the ProQuest Thesaurus
While LCSH serves as the primary source for the CDI-normalized subjects, MeSH headings are used to add additional subjects and alternate terms that do not exist in LCSH. CDI uses the mapping between LCSH and MeSH from Northwestern University Libraries, which can be found at https://galter.northwestern.edu/about-us/northwestern-university-libraries-lcsh-mesh-mapping-project
MeSH subjects and their entry terms are added as alternate terms for the corresponding LCSH subject if there is a mapping entry in the Northwestern University Libraries file. If they cannot be mapped to an LCSH subject, MeSH terms are added as new subjects to the CDI list of subjects. Their entry terms are added as the alternate terms for the subject unless they are programmatically determined to be mappings to narrower subjects, mappings to broader subjects, or ambiguous subjects. For example, "Calcimycin" (D000001) is added as a subject, and "Antibiotic A23187" is added as its alternate term.
In addition, we use the ProQuest Thesaurus to include a small set of additional alternate terms that closely match the LCSH normalized subjects.

