Home
WordLevel is an open-source library for word-level natural language processing and text analysis. It allows you to measure how two bodies of text differ token by token, and turn that measurement into a publication-ready word shift graph: a chart that shows not only which words distinguish the texts, but also how they do so.
Install¶
WordLevel runs on Python 3.10 and above, and installs with any Python package manager. See the install guide for details.
Compare texts¶
The code below compares the speeches of Presidents Franklin D. Roosevelt and Joe Biden by sentiment, scoring each president with a lexicon from each era.
import wordlevel as wl
speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Franklin D. Roosevelt", "Joe Biden"])
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.lexicon(
lexicon_reference=wl.lex.SocialSent("1940"),
lexicon_comparison=wl.lex.SocialSent("2000"),
reference_score="center",
)
.normalize(by="total_diff")
.alias("roosevelt_biden"),
)
chart = cl.plot.shift("roosevelt_biden", show_totals=True)
Biden's speeches score lower in sentiment overall. From the word shift graph, we can see that most of the gap comes from words like "nation," "families," and "americans" carrying lower sentiment in the 2000s lexicon than the 1940s one (▽), only partly offset by Biden using relatively positive words such as "together" and "well" more often (+↑).
Where to go next¶
-
Quick Start
Install the package, make your first comparison, and build a word shift graph.
-
Cookbooks
Task-by-task recipes for constructing catalogs, scoring, and plotting.
-
Concepts
The ideas behind corpora, comparisons, and word shift graphs.
-
API Reference
Every public class, method, and configuration, generated from the source.
Citation¶
If you use WordLevel in your research, please cite the following paper:
Gallagher, R. J., Frank, M. R., Mitchell, L., Schwartz, A. J., Reagan, A. J., Danforth, C. M., & Dodds, P. S. (2021). Generalized word shift graphs: a method for visualizing and explaining pairwise comparisons between texts. EPJ Data Science, 10(1), 4.