AI & Language · September 2026 · 7 min read
Last week I asked the same AI question twice. Once in Portuguese. Once in English. The answers weren’t identical — and that difference says a lot about how today’s AI really works.
I live in Brazil. I write about AI in Portuguese. And every single day, I use tools that were built — in their bones, in their training data, in their default assumptions — for someone who isn’t me. That’s not a complaint. It’s just the reality of using AI in 2026 when your first language isn’t English. And once you understand that reality clearly, you can actually work with it — instead of being silently limited by it.
This post is for the 95% of the world that Silicon Valley tends to forget.
“English dominates much of the publicly available text used to train modern AI models — while many other languages remain significantly underrepresented.”
Tilde AI / LLM Training Data Research, 2026
First, understand what you’re actually dealing with
The large language models powering ChatGPT, Claude, Gemini and the rest weren’t trained equally across languages. They were fed mostly English text — news articles, books, Reddit threads, academic papers, Stack Overflow answers. Portuguese, Spanish, Hindi, Arabic — all present, but in much smaller fractions.
The LLaMA model, developed by Meta and used as a foundation for dozens of AI tools, was trained on 2 trillion tokens of text. Of those, only 2.8 billion were in Portuguese — 0.09% of the total. That number alone explains a lot.
What this means in practice: when you write to the AI in another language, it responds — often well. But the cultural references it reaches for, the examples it defaults to, the assumptions buried in its answers — all of them tend to reflect an English-speaking world. Knowing this changes how you use the tool.
The strategies that actually work
1. For technical tasks, consider prompting in English — but don’t assume you always need to
For highly technical topics — coding, science, legal analysis, academic writing — English prompts often produce more precise or detailed responses. This happens because much of the technical documentation and specialized knowledge in AI training data is in English.
But the gap has narrowed significantly with newer models. For everyday writing, content creation, or general questions, prompting in your own language works very well. Test both. The difference isn’t always what you’d expect — and it varies by model and task.
2. Give context the AI can’t assume — this is the most important one
The AI doesn’t automatically know what country you’re in, what laws apply to you, or what your audience expects. You have to tell it — and the more specific you are, the better the result.
Instead of: “Write an article.”
Try: “Write a 1,500-word article for Brazilian professionals, in accessible language, with examples from the public sector, optimized for SEO.”
Context isn’t optional. It’s the whole game.
3. Use English for research, your language for creation
When I’m researching a topic — understanding concepts, finding angles, exploring ideas — I often do it in English. New AI research, technical documentation, and global news tend to appear in English first, sometimes months before a quality version exists in other languages.
Then, when I’m ready to write or create something, I switch to Portuguese. This hybrid workflow sounds counterintuitive but it cuts research time significantly and produces better raw material to work with.
4. Always verify local facts independently
AI hallucinates in every language — but it does so more confidently when it has less data to draw from. Local laws, municipal regulations, regional prices, internal procedures, recent local news: always cross-check these against a local source.
I treat AI-generated local facts the same way I’d treat a tip from a well-meaning friend who has never visited my country. Useful starting point. Not a final answer.
5. Use AI to sound natural in English — not just correct
One of the most underrated uses of AI for non-English speakers isn’t translation — it’s tone calibration. When I write something in English for an international audience, I write a rough version first and then ask: “Rewrite this so it sounds natural to a native English speaker — keep my ideas exactly but fix anything that sounds translated.”
It catches things no dictionary would catch: expressions that are technically correct but feel foreign, sentence structures that reveal a non-native writer, idioms that don’t cross cultural lines cleanly.
6. Pick the right model for your language
Not all AI models handle non-English languages equally. In my daily experience writing in Portuguese, Claude and ChatGPT handle Brazilian Portuguese best for nuanced writing tasks. Gemini performs well for factual queries. Some open-source models still struggle with regional slang and informal registers.
Test them yourself with a task specific to your language and context. Don’t assume the most popular model is the best one for your needs.
The bigger picture — and why it matters beyond productivity
These strategies make you more effective today. But there’s a larger issue worth naming.
When AI is trained primarily on English data, it doesn’t just perform differently in other languages — it carries an English-speaking worldview into every interaction. The assumptions baked into its answers, the examples it reaches for, the cultural norms it treats as default — all of them reflect a specific slice of humanity, not humanity as a whole.
Brazil alone has 215 million people. The Portuguese-speaking world has over 260 million. Add Hindi, Arabic, Swahili, Bengali — and you’re talking about billions of people whose languages, cultures and daily realities are underrepresented in tools that are rapidly reshaping the world.
What “underrepresented in AI” actually looks like
- —Medical or legal responses that miss regional laws and local practices
- —Chatbots that don’t understand regional accents or informal speech patterns
- —Educational tools built around American or European cultural references
- —AI-generated content that sounds subtly foreign even when grammatically correct
- —Voice recognition that performs worse for non-English speakers
Until the language gap closes, the practical answer is this: consider English for highly technical tasks, give the AI context it can’t infer, verify anything local or time-sensitive, and take advantage of being bilingual. Accessing knowledge in both English and your native language lets you combine the latest global information with local understanding — something many users simply can’t do. That’s not a small thing. That’s a real edge.
Tags:
AI Tips · Non-English AI · Language Gap · Portuguese · AI Prompting · Global South · ChatGPT · Claude