BERTomelo: Your Portuguese Encoder Best Friend
- lab arXivLabs
- location Portugal
- person Luís Paulo Faina Garcia
- product Albertina
- product BERTimbau
- product BERTomelo
- product ClassiCC-PT
- product FlashAttention
A team of researchers has released BERTomelo, a monolingual encoder for Portuguese built on the ModernBERT architecture and trained from scratch on a corpus of 106 million documents, according to a paper published on arXiv [1]. The model, introduced by authors including Luís Paulo Faina Garcia, is designed to address what the researchers describe as a gap in Portuguese natural language processing, where existing monolingual encoders such as BERTimbau and Albertina have lagged behind English benchmarks in scalability and efficiency [1][2]. BERTomelo is offered in Base and Large versions, both featuring a 1,024-token context window and hardware-level optimizations including FlashAttention and alternating attention mechanisms [1][4]. The model was pretrained on ClassiCC-PT, a high-quality Portuguese corpus that the authors say ensures alignment with contemporary usage of the language [1][2]. The Base variant contains 136 million parameters across 22 layers, while the Large variant holds 377 million parameters across 28 layers, according to model cards published on Hugging Face [5][6]. Pretraining encompassed 386.48 billion tokens, and the custom tokenizer was sourced from ModBERTBr and trained using a Unigram algorithm on the BrWaC corpus and Portuguese Wikipedia [5][6]. In benchmark evaluations, BERTomelo Large achieved a Named Entity Recognition F1 score of 91.05 percent, with recall at 92.75 percent and precision at 89.42 percent, outperforming mBERT, BERTimbau Large, and Albertina 900M on that task [5][6]. On Recognizing Textual Entailment, the Large model reached an F1 of 89.21 percent and accuracy of 89.26 percent [6]. For Semantic Textual Similarity, it recorded a mean squared error of 0.401 and a Pearson correlation of 0.849 [6]. The release comes as the Portuguese NLP community continues to build evaluation infrastructure. A separate initiative, the CLARIN-PT-LDB leaderboard, was recently launched to assess open large language models on European Portuguese using ten benchmarks, including novel tests for cultural alignment and model safeguards [9]. The BERTomelo authors note that standardized long-context benchmarks for Portuguese remain unavailable and that the current release does not yet incorporate a sequence packing mechanism [5][6].
tool-releaseresearch-paperapplication
Background sources we checked (10)
- arxiv.org ↗ Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models offer broad coverage, dedicated monolingual encoders are essential for capturing the unique lexical and syntactic nuances of speci…
- arxiv.org ↗ [2606.28999] BERTomelo: Your Portuguese Encoder Best Friend ... # Title:BERTomelo: Your Portuguese Encoder Best Friend ... Authors: Rennê Ruan Alves Oliveira, Gustavo Cordeiro Galvão Van Erven, Luís Paulo Faina Garcia ... > Abstract:Encoders have become the state of the art for m…
- arxiv.org ↗ ## BERTOMELO: YOUR PORTUGUESE ENCODER BEST FRIEND ... ABSTRACT Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models offer broad coverage, dedicated monolingual encoders are essential …
- huggingface.co ↗ BERTomelo is a variant of BERT encoders built upon the ModernBERT architecture. It is pretrained from scratch on the Classified Common Crawl Corpus for Portuguese (ClassiCC PT), specifically tailored and specialized for the Portuguese language. BERTomelo try to fill the gap of ol…
- huggingface.co ↗ # BERTomelo ... BERTomelo is a variant of BERT encoders built upon the ModernBERT architecture. It is pretrained from scratch on the Classified Common Crawl Corpus for Portuguese (ClassiCC PT), specifically tailored and specialized for the Portuguese language. BERTomelo try to fi…
- arxiv.org ↗ [2603.26511] AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese ... # Title:AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese ... > Abstract:Despite rapid progress in open large language models (LLMs),…
- arxiv.org ↗ Authors: Pedro Quaresma(CISUC / Department of Mathematics, University of Coimbra, Portugal), Vanda Santos(CIDTFF / University of Aveiro and CISUC, Portugal) ... About arXivLabs ... # arXivLabs: experimental projects with community collaborators ... arXivLabs is a framework that a…
- arxiv.org ↗ -PT-L ... # CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility ... This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leader…
- en.wikipedia.org ↗ Carlos Ernesto Guestrin (born 1975) is a Brazilian computer scientist and a professor at Stanford University. He is best known for his contributions to scalable machine learning algorithms.…
- en.wikipedia.org ↗ Vladlen Koltun (Hebrew: ולדלן קולטן; born 1980) is an Israeli-American computer scientist and intelligent systems researcher. He currently serves as distinguished scientist at Apple Inc. His main areas of research are artificial intelligence, computer vision, machine learning, an…
Sources covering this (2)
- export.arxiv.org — BERTomelo: Your Portuguese Encoder Best Friend ↗
- export.arxiv.org — Friend or Foe · Global