📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal launched AMÁLIA, a European Portuguese large language model, with promising benchmarks. However, experts raise three fundamental questions about its openness, data sufficiency, and objectives. These questions have broader implications for Europe’s sovereign AI efforts.
Portugal’s €5.5 million AMÁLIA large language model is now operational, marking a significant step in the country’s AI efforts. Developed by a consortium of research institutions, it outperforms previous open models on Portuguese benchmarks but faces scrutiny over fundamental structural questions about its openness, data, and objectives, which are critical for national AI policy.
AMÁLIA was announced in December 2024 and completed its base version by September 2025, with a public launch in October 2025. It is built as a continuation of the EuroLLM multilingual foundation, involving around 60 researchers across Portugal’s top research institutions. The model is currently used by 450,000 academic users and holds knowledge up to the end of 2023. Its benchmarks show it surpasses previous open models on most Portuguese tasks, though it still trails Qwen 3-8B on some benchmarks like ALBA.
Despite these technical achievements, prominent analysts, including Duarte O.Carmo, have raised three core questions: How open is ‘fully open’ in practice? How much native-language data is enough? And what should be the primary goal of such models? These questions are not yet answered publicly by the Portuguese team or other European projects, raising concerns about transparency and strategic direction.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.

Advanced Language Tool Kit: Teaching the Structure of the English Language
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.

ESSENTIAL AI TOOLS FOR TRANSPARENT MODELS USING SHAP, LIME, AND VISUALIZATION TECHNIQUES: 65 PRACTICAL EXERCISES TO ENHANCE INTERPRETABILITY AND TRUST IN BLACK-BOX MODELS
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.
multilingual AI training datasets
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.

AI & Machine Learning: Generative AI, LLMs & AI Ethics — European Edition (Volume 3): Advanced Guide to Large Language Models, Prompt Engineering, and Responsible AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications for Portugal’s AI Strategy and European Sovereignty
The questions surrounding AMÁLIA reveal broader issues facing Europe’s sovereign AI initiatives, such as transparency, data sufficiency, and goal-setting. The answers will influence how European countries develop and govern their AI models, impacting national security, competitiveness, and digital sovereignty. The model’s current state underscores the need for clear strategic frameworks to guide future investments and research priorities.
European Sovereign-Language Model Development Landscape
Across Europe, several countries are developing their own large language models, including Italy’s Minerva, Germany’s Aleph Alpha, and France’s Mistral, often with similar structural questions. The European Union and various national governments have committed significant funds, emphasizing sovereignty and data control. However, the discourse often focuses on individual models rather than the systemic issues, such as openness, data sufficiency, and purpose, which are critical to the success of these initiatives.
Portugal’s AMÁLIA exemplifies this pattern: a publicly funded project with promising benchmarks but facing fundamental questions about its design and strategic goals. The broader European effort remains opaque on these core issues, risking fragmented development and suboptimal outcomes.
“The three questions—openness, data sufficiency, and objectives—are essential to understanding what these models can truly achieve.”
— Duarte O.Carmo
Unanswered Questions About AMÁLIA’s Openness, Data, and Goals
It remains unclear how open AMÁLIA truly is in practice, especially regarding access and transparency. The sufficiency of native Portuguese data used during training, particularly the 5.8 billion tokens from Arquivo.pt, is also debated. Additionally, the primary objectives—whether to prioritize performance, transparency, or strategic sovereignty—have not been publicly clarified by the project team. The final version, due in June 2026, may address some of these gaps, but current details are limited.
Upcoming Milestones and Strategic Clarifications
The final version of AMÁLIA is scheduled for release in June 2026, which will likely include further technical improvements and possibly more transparency. Over the next 12 to 24 months, European projects are expected to face increased scrutiny regarding their openness, data policies, and strategic objectives. Policymakers and researchers will need to address these foundational questions to ensure the models serve broader national and regional interests effectively.
Key Questions
What makes AMÁLIA different from other European language models?
AMÁLIA is built as a continuation of a multilingual foundation, rather than from scratch, and is publicly funded by Portugal. It is designed specifically for European Portuguese and involves a large academic consortium, making it a key case in Europe’s sovereign AI efforts.
Why are the questions of openness, data, and goals so important?
These questions determine the transparency, strategic direction, and long-term viability of national AI models. Addressing them is essential for ensuring models are trustworthy, aligned with national interests, and capable of competing globally.
What risks does the current uncertainty pose for Portugal and Europe?
Without clear answers, there is a risk of fragmented development, lack of trust, and suboptimal use of public funds. It could also hinder Europe’s ability to develop sovereign AI that aligns with regional values and security needs.
Will the final version of AMÁLIA resolve these questions?
It is not yet certain. The final version may address some concerns, but the core structural questions about openness, data sufficiency, and objectives require explicit strategic decisions that may extend beyond technical improvements.
Source: ThorstenMeyerAI.com