Wikidata Lexemes Scraper avatar

Wikidata Lexemes Scraper

Pricing

from $10.00 / 1,000 result items

Go to Apify Store
Wikidata Lexemes Scraper

Wikidata Lexemes Scraper

Search and extract Wikidata Lexemes (L-namespace). Returns lemma, language QID, lexical category, senses, glosses, statements, and optional inflected forms for each lexeme. Distinct from Q-entities.

Pricing

from $10.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

ParseForge Banner

🧬 Wikidata Lexemes Scraper

πŸš€ Export structured lexicographic data in seconds. Pull lemmas, lexical categories, grammatical forms, and senses from the Wikidata Lexeme namespace across hundreds of languages. No API key, no registration, no SPARQL skills required.

The Wikidata Lexemes Scraper queries the L-namespace on wikidata.org and returns 14 fields per record, including lexeme ID, lemma, language QID, lexical category QID, every documented sense, every inflected form, statement metadata, and a link back to the canonical Wikidata Lexeme page. The L-namespace is a structured, machine-readable companion to Wiktionary and powers downstream dictionaries, linguistic research, and language documentation.

The dataset covers more than 1.4 million lexemes spanning over 1,500 languages, from major world languages down to documented endangered and historical languages. This Actor turns the namespace into downloadable CSV, Excel, JSON, or XML in under five minutes. Lemma search, language filter, and lexical category filter all run from the same input form.

🎯 Target AudienceπŸ’‘ Primary Use Cases
Linguists, NLP engineers, dictionary builders, language-documentation teams, computational lexicographers, knowledge-graph engineersMultilingual lemma dictionaries, training data for morphology models, inflection tables for language apps, structured glosses for translation pipelines

πŸ“‹ What the Wikidata Lexemes Scraper does

Three lookup workflows in a single run:

  • πŸ” Lemma search. Query the L-namespace by any lemma string in any UI language.
  • 🌐 Language filter. Restrict results to a single language QID such as Q1860 English or Q150 French.
  • πŸ”€ Lexical category filter. Limit to a part of speech via QID, for example Q1084 noun or Q24905 verb.

Each record includes the lexeme ID, lemma, lemma language code, language QID, lexical category QID, a short description, every documented sense with multilingual glosses, every inflected form with grammatical feature QIDs, statement count and properties, last-modified timestamp, the canonical Lexeme URL, and the scrape timestamp.

πŸ’‘ Why it matters: structured lexicographic data powers morphological analyzers, inflection tables, multilingual search, and machine translation. Building your own pipeline means writing SPARQL queries, handling pagination across the L-namespace, and joining sense and form data by hand. This Actor skips all of that and refreshes on every run.

πŸ“Š Data fields

Each record includes: description, formCount, forms, languageQid, lastModified, lemma, lemmaLanguage, lexemeId, lexemeUrl, lexicalCategoryQid, scrapedAt, senseCount, senses, statementCount, statementProperties. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

πŸš€ How to use

  1. πŸ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the Wikidata Lexemes Scraper page on the Apify Store.
  3. 🎯 Set input. Enter a lemma search, pick a language QID and lexical category QID, and set maxItems.
  4. πŸš€ Run it. Click Start and let the Actor collect your data.
  5. πŸ“₯ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Wikidata, the Wikimedia Foundation, or any of its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open lexicographic data is collected.

πŸ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.