Wikidata Lexemes Scraper
Pricing
from $10.00 / 1,000 result items
Wikidata Lexemes Scraper
Search and extract Wikidata Lexemes (L-namespace). Returns lemma, language QID, lexical category, senses, glosses, statements, and optional inflected forms for each lexeme. Distinct from Q-entities.
Pricing
from $10.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share

𧬠Wikidata Lexemes Scraper
π Export structured lexicographic data in seconds. Pull lemmas, lexical categories, grammatical forms, and senses from the Wikidata Lexeme namespace across hundreds of languages. No API key, no registration, no SPARQL skills required.
The Wikidata Lexemes Scraper queries the L-namespace on wikidata.org and returns 14 fields per record, including lexeme ID, lemma, language QID, lexical category QID, every documented sense, every inflected form, statement metadata, and a link back to the canonical Wikidata Lexeme page. The L-namespace is a structured, machine-readable companion to Wiktionary and powers downstream dictionaries, linguistic research, and language documentation.
The dataset covers more than 1.4 million lexemes spanning over 1,500 languages, from major world languages down to documented endangered and historical languages. This Actor turns the namespace into downloadable CSV, Excel, JSON, or XML in under five minutes. Lemma search, language filter, and lexical category filter all run from the same input form.
| π― Target Audience | π‘ Primary Use Cases |
|---|---|
| Linguists, NLP engineers, dictionary builders, language-documentation teams, computational lexicographers, knowledge-graph engineers | Multilingual lemma dictionaries, training data for morphology models, inflection tables for language apps, structured glosses for translation pipelines |
π What the Wikidata Lexemes Scraper does
Three lookup workflows in a single run:
- π Lemma search. Query the L-namespace by any lemma string in any UI language.
- π Language filter. Restrict results to a single language QID such as
Q1860English orQ150French. - π€ Lexical category filter. Limit to a part of speech via QID, for example
Q1084noun orQ24905verb.
Each record includes the lexeme ID, lemma, lemma language code, language QID, lexical category QID, a short description, every documented sense with multilingual glosses, every inflected form with grammatical feature QIDs, statement count and properties, last-modified timestamp, the canonical Lexeme URL, and the scrape timestamp.
π‘ Why it matters: structured lexicographic data powers morphological analyzers, inflection tables, multilingual search, and machine translation. Building your own pipeline means writing SPARQL queries, handling pagination across the L-namespace, and joining sense and form data by hand. This Actor skips all of that and refreshes on every run.
π Data fields
Each record includes: description, formCount, forms, languageQid, lastModified, lemma, lemmaLanguage, lexemeId, lexemeUrl, lexicalCategoryQid, scrapedAt, senseCount, senses, statementCount, statementProperties. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
π How to use
- π Sign up. Create a free account with $5 credit (takes 2 minutes).
- π Open the Actor. Go to the Wikidata Lexemes Scraper page on the Apify Store.
- π― Set input. Enter a lemma search, pick a language QID and lexical category QID, and set
maxItems. - π Run it. Click Start and let the Actor collect your data.
- π₯ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.
β±οΈ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
π Recommended Actors
- π Wiktionary Definitions Scraper - Multilingual dictionary entries with definitions and examples
- π Wikipedia Scraper - Encyclopedic articles and references
- π° ArXiv Scraper - Scientific preprint metadata
- π Indexmundi Scraper - Global demographic and economic indicators
- πΊοΈ Nominatim OSM Scraper - Geocode addresses via OpenStreetMap
π‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
β οΈ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Wikidata, the Wikimedia Foundation, or any of its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open lexicographic data is collected.
π Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.