4  Language

MARC: 041,a,h, 008

4.1 Description

The polish_languages function is designed to standardize and harmonize language information in a dataset. The process starts by isolating unique language entries, ensuring that each distinct combination of languages is processed only once, which improves efficiency and avoids redundant computations. A MARC reference list of recognized language abbreviations and names is then used to map language codes to their standardized forms. Each language entry is analyzed to identify multiple languages and to detect any unrecognized terms.

The entries are standardized by converting them to their recognized forms while eliminating duplicates and filtering out unrecognized languages. Empty cells in the dataset are marked as NA to indicate missing information. Once standardized, all valid languages are aggregated to create a structured data frame. This data frame includes the total number of languages in each entry, a flag (TRUE/FALSE) indicating whether the entry contains multiple languages (including those that are originally coded as mul = Multiple language), the cleaned and harmonized list of languages, and the primary language, which is defined as the first listed language in each entry. The result is a cleaned and standardized dataset that facilitates accurate analysis of multilingual data.

Additionally, an error list is generated, consisting of unrecognized language information and the corresponding IDs. This error list helps librarians identify mistakes in the original data and provides context to either correct the errors or explain why certain entries were discarded by the function.

This page reports two related outputs. The first section combines publication-language information from MARC 041$a and 008. Field 041$a is used as the main source, and missing 041$a values are filled from 008. The second section reports original-language information from MARC 041$h, which is used as translation-related metadata. Error lists are produced separately for unrecognized values so that problematic records can be reviewed by librarians.

4.2 Field 041$a and 008 combined

This section reports harmonized publication-language information. The variable is based primarily on MARC 041$a, while missing 041$a values are filled from MARC 008. The resulting combined field is therefore intended as the main publication-language variable for record-level analysis. There are 0 records with missing author name information in the original raw data.

4.2.1 Complete Dataset Overview

Unique languages: 1.04537^{5}

Unique primary languages: 142

1307778 single-language entries (99.25%)

9879 multilingual entries , accounting for 0.75% of the total. This includes entries explicitly coded as “mul” (Multiple languages) as well as those with more than one language listed for a single book.

There are 230 single-language entries marked as only “Undetermined”, coded as “und”, accounting for (0.02%) of the total.

There are 406 missing values in the dataset,accounting for (0.03%) of the total.

Conversions from raw to preprocessed language entries

Download language harmonized dataset

Language Entries (n) Fraction (%)
Finnish 599941 45.5
Unrecognized 517244 39.2
Swedish 94981 7.2
English 69789 5.3
Multiple languages 8431 0.6
German 7715 0.6
Russian 4535 0.3
Latin 3722 0.3
French 2639 0.2
Unrecognized;Unrecognized;Unrecognized 1330 0.1

Unrecognized languages 041a provides details of languages that were discarded, in total: 5. Additionally, the Error list 041a contains ID numbers of entries associated with these discarded languages, intended for librarian review.

Unrecognized languages 008 provides details of languages that were discarded, in total: 16. Additionally, the Error list 008 contains ID numbers of entries associated with these discarded languages, intended for librarian review.

4.2.2 Subset Analysis: 1809-1917

Unique languages (1809-1917): 4731

Unique primary languages (1809-1917): 35

72346 single-language entries (99.61%)

280 multilingual entries , accounting for 0.39% of the total. This includes entries explicitly coded as “mul” (Multiple languages) as well as those with more than one language listed for a single book.

There are 9 entries marked as “Undetermined”.

There are 0 missing values in the dataset,accounting for (0%) of the total.

Download language harmonized dataset (1809-1917)

4.2.3 Top languages for 1809-1917

Number of titles assigned with each language (top-10). For a complete list, see accepted languages (1809-1917).

Language Entries (n) Fraction (%)
Unrecognized 35084 48.3
Finnish 21782 30
Swedish 12809 17.6
German 849 1.2
Russian 539 0.7
Latin 378 0.5
French 326 0.4
Multiple languages 264 0.4
English 256 0.4
Estonian 62 0.1

Title count per language (including multi-language documents; note the log10 scale):

4.3 Field 041,h: Original Language of Translations

MARC field 041 subfield h records the original language of a work when the bibliographic record describes a translation.

This field should be interpreted as translation-related metadata: records with a value in language_original are records for which Fennica contains information about the original language of the work. However, the absence of 041,h does not necessarily prove that the record is not a translation; it only means that no original-language information was recorded or extracted from this field. There are 0 records with missing author name information in the original raw data.

4.3.1 Complete Dataset Overview

Unique original languages: 190

Unique primary original languages: 0

There are 1256548 records with original-language information in 041$h, accounting for 95.33% of the complete dataset. These records can be treated as translation records with recorded original-language metadata.

There are 61515 records without original-language information in 041\(h, accounting for 4.67% of the complete dataset. These records should not automatically be interpreted as non-translations; they only lack recorded 041\)h information.

1178773 records have one original language recorded. This corresponds to 93.81% of records with 041$h information.

77775 records have multiple original languages recorded, accounting for 6.19% of records with 041$h information. This includes entries explicitly coded as “Multiple languages” as well as records where more than one original language is listed.

There are 0 records where the original language is marked as “Undetermined”, coded as “und”. This accounts for 0% of records with 041$h information.

Unrecognized original languages provides details of language values that were discarded during harmonization, in total: 18.

Additionally, the error list contains record IDs associated with discarded original-language values. These records are intended for librarian review.

Conversions from raw to preprocessed original-language entries

Original language Entries (n) Fraction of records with 041$h (%)
fin 842715 67.1
eng 145346 11.6
swe 134191 10.7
fin;swe 28305 2.3
ger 14883 1.2
lat 14155 1.1
fin;eng 12728 1
fin;swe;eng 7376 0.6
rus 5699 0.5
mul 5028 0.4

Download original-language harmonized dataset

4.3.2 Subset Analysis: 1809–1917

Unique original languages, 1809–1917: 55

Unique primary original languages, 1809–1917: 0

There are 71296 records with original-language information in 041$h in the 1809–1917 subset, accounting for 98.17% of the subset.

There are 1330 records without original-language information in 041$h, accounting for 1.83% of the 1809–1917 subset.

67453 records have one original language recorded. This corresponds to 94.61% of records with 041$h information in the 1809–1917 subset.

3843 records have multiple original languages recorded, accounting for 5.39% of records with 041$h information.

There are 0 records where the original language is marked as “Undetermined”.

Download original-language harmonized dataset, 1809–1917

4.3.3 Top Original Languages for 1809–1917

The following table shows the most common original languages recorded in 041$h for the 1809–1917 subset. Counts include records with multiple original languages. For a complete list, see accepted original languages, 1809–1917.

Original language Entries (n) Fraction of records with 041$h (%)
fin 37707 52.9
swe 22270 31.2
lat 2506 3.5
ger 2283 3.2
fin;swe 1740 2.4
rus 848 1.2
fre 702 1
eng 452 0.6
swe;fin 252 0.4
swe;lat 203 0.3

Title count per original language, including records with multiple original languages. Note the log10 scale: