Tatoeba Sentence Corpus Scraper
Pricing
from $10.00 / 1,000 result items
Tatoeba Sentence Corpus Scraper
Extract Tatoeba sentence corpus with millions of bilingual example sentences. Capture sentence ID, language, text, owner, audio URL, translations, tags, and license. Export to JSON, CSV, or Excel for language learning, NLP training data, translation memory, and linguistic research.
Pricing
from $10.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share

π£οΈ Tatoeba Sentence Corpus Scraper
π Export the world's largest open multilingual sentence corpus in seconds. Pull 12,000,000+ example sentences across 400+ languages with translations, audio, contributor info, and CC-BY licence metadata. No login, no manual CSV stitching.
The Tatoeba Sentence Corpus Scraper taps into the Tatoeba community catalog and returns 14 structured fields per sentence, including the original text, language code and name, translation list, audio links, contributor handle, correctness score, and licence. Tatoeba has been collaboratively edited by linguists, polyglots, and language learners since 2006, and ships under a permissive Creative Commons licence.
The catalog covers every major living language family plus dozens of constructed, classical, and minority languages, from Mandarin and Spanish down to Latin, Esperanto, and revived regional tongues. This Actor turns that into a clean CSV, Excel, JSON, or XML dataset in under five minutes, with all filtering done server-side so you skip the parsing entirely.
| π― Target Audience | π‘ Primary Use Cases |
|---|---|
| Linguists, language-learning app builders, translation researchers, NLP engineers, lexicographers, ESL teachers, audio dataset curators | Parallel corpus mining, flashcard sourcing, translation memory seeding, speech model training, idiom and proverb research, classroom example banks |
π What the Tatoeba Scraper does
Four sentence-mining workflows in a single run:
- π Keyword search. Find every sentence containing a target word or phrase.
- π Source language filter. Pick a single source language out of 400+ (Tatoeba uses ISO 639-3 codes).
- βοΈ Target translation filter. Restrict to sentences that have a translation into a chosen language.
- π·οΈ Tag filter. Pull only sentences tagged with concepts like
proverb,idiom,greeting, or any community label.
Each record includes the sentence ID, raw text, language code and human-readable name, every linked translation (with that translation's language), audio file URLs when available, contributor handle, correctness score, licence, and the canonical Tatoeba page link.
π‘ Why it matters: clean, licence-clear parallel sentences are the raw material of every translation memory, language-learning flashcard, and speech model training set. Building your own pipeline against the Tatoeba site means writing fragile HTML parsers and respecting rate limits by hand. This Actor delivers the same data structured and ready to import.
π Data fields
Each record includes: audioUrls, contributor, correctness, direction, hasAudio, language, languageName, license, scrapedAt, sentenceId, text, translationCount, translations, url. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
π How to use
- π Sign up. Create a free account with $5 credit (takes 2 minutes).
- π Open the Actor. Go to the Tatoeba Sentence Corpus Scraper page on the Apify Store.
- π― Set input. Pick a source language, optional keyword, optional target language and tags, set
maxItems. - π Run it. Click Start and let the Actor collect your sentences.
- π₯ Download. Grab your results from the Dataset tab as CSV, Excel, JSON, or XML.
β±οΈ Total time from signup to downloaded corpus: 3-5 minutes. No coding required.
π Recommended Actors
- π MyMemory Translation Scraper - Translate text across 70+ language pairs
- π LibriVox Audiobooks Scraper - Public-domain audiobooks with reader credits
- ποΈ Library of Congress Scraper - 170M+ digitized cultural records
- π° ArXiv Scraper - Academic preprints with metadata
- π Figshare Scraper - Open research datasets and figures
π‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
β οΈ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Tatoeba Project or its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open corpus data is collected, under the project's Creative Commons licence.
π Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.
