Tatoeba Sentence Corpus Scraper avatar

Tatoeba Sentence Corpus Scraper

Pricing

from $10.00 / 1,000 result items

Go to Apify Store
Tatoeba Sentence Corpus Scraper

Tatoeba Sentence Corpus Scraper

Extract Tatoeba sentence corpus with millions of bilingual example sentences. Capture sentence ID, language, text, owner, audio URL, translations, tags, and license. Export to JSON, CSV, or Excel for language learning, NLP training data, translation memory, and linguistic research.

Pricing

from $10.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Share

ParseForge Banner

πŸ—£οΈ Tatoeba Sentence Corpus Scraper

πŸš€ Export the world's largest open multilingual sentence corpus in seconds. Pull 12,000,000+ example sentences across 400+ languages with translations, audio, contributor info, and CC-BY licence metadata. No login, no manual CSV stitching.

The Tatoeba Sentence Corpus Scraper taps into the Tatoeba community catalog and returns 14 structured fields per sentence, including the original text, language code and name, translation list, audio links, contributor handle, correctness score, and licence. Tatoeba has been collaboratively edited by linguists, polyglots, and language learners since 2006, and ships under a permissive Creative Commons licence.

The catalog covers every major living language family plus dozens of constructed, classical, and minority languages, from Mandarin and Spanish down to Latin, Esperanto, and revived regional tongues. This Actor turns that into a clean CSV, Excel, JSON, or XML dataset in under five minutes, with all filtering done server-side so you skip the parsing entirely.

🎯 Target AudienceπŸ’‘ Primary Use Cases
Linguists, language-learning app builders, translation researchers, NLP engineers, lexicographers, ESL teachers, audio dataset curatorsParallel corpus mining, flashcard sourcing, translation memory seeding, speech model training, idiom and proverb research, classroom example banks

πŸ“‹ What the Tatoeba Scraper does

Four sentence-mining workflows in a single run:

  • πŸ”Ž Keyword search. Find every sentence containing a target word or phrase.
  • 🌐 Source language filter. Pick a single source language out of 400+ (Tatoeba uses ISO 639-3 codes).
  • ↔️ Target translation filter. Restrict to sentences that have a translation into a chosen language.
  • 🏷️ Tag filter. Pull only sentences tagged with concepts like proverb, idiom, greeting, or any community label.

Each record includes the sentence ID, raw text, language code and human-readable name, every linked translation (with that translation's language), audio file URLs when available, contributor handle, correctness score, licence, and the canonical Tatoeba page link.

πŸ’‘ Why it matters: clean, licence-clear parallel sentences are the raw material of every translation memory, language-learning flashcard, and speech model training set. Building your own pipeline against the Tatoeba site means writing fragile HTML parsers and respecting rate limits by hand. This Actor delivers the same data structured and ready to import.

πŸ“Š Data fields

Each record includes: audioUrls, contributor, correctness, direction, hasAudio, language, languageName, license, scrapedAt, sentenceId, text, translationCount, translations, url. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

πŸš€ How to use

  1. πŸ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the Tatoeba Sentence Corpus Scraper page on the Apify Store.
  3. 🎯 Set input. Pick a source language, optional keyword, optional target language and tags, set maxItems.
  4. πŸš€ Run it. Click Start and let the Actor collect your sentences.
  5. πŸ“₯ Download. Grab your results from the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to downloaded corpus: 3-5 minutes. No coding required.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Tatoeba Project or its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open corpus data is collected, under the project's Creative Commons licence.

πŸ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.