AutoSpecNER: A Fine-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction
- lab arXiv
- lab arXivLabs
- person Filippos Ventirozos
Researchers have released AutoSpecNER, an expert-annotated dataset designed to improve fine-grained entity recognition in vehicle advertisements, according to a paper posted on arXiv [1]. The dataset addresses a gap in automotive natural language processing resources [1]. The dataset comprises 659 advertisements sourced from a popular car-selling website, containing over 10,000 annotated entities across 15 distinct categories such as MODEL, ENGINE_SPEC, and BATTERY_CAPACITY [1]. The annotation process was validated through inter-annotator agreement, which reached an average score of 91.5% [1]. The paper was submitted to arXiv, an open-access repository for electronic preprints that has hosted over two million articles since its founding in 1991, on June 23, 2026, by Filippos Ventirozos [6][1]. To establish performance baselines, the researchers benchmarked several approaches. A rule-based extraction system achieved a micro-F1 score of 43% [1]. The strongest large language model tested reached 77.8% [1]. Large language models are machine learning systems with many parameters trained on vast amounts of text for natural language processing tasks [8]. The fine-tuned transformer encoder DeBERTa outperformed both, recording a 90% micro-F1 score [1]. The paper appears on arXiv, which as of late 2024 was receiving approximately 24,000 new submissions per month [6]. The platform also supports community-developed tools through its arXivLabs framework, a formalized collaboration space launched to allow third parties to build features that enhance the reading experience while adhering to values of openness and user data privacy [4]. These experimental projects, accessible via tabs on the abstract page, include tools such as the Bibliographic Explorer for navigating citation trees and the CORE Recommender for discovering related open-access papers [5][4].
research-paperbenchmark
Background sources we checked (7)
- arxiv.org ↗ Vehicle advertisements contain rich specification information, but automotive NER resources remain limited. We introduce AutoSpecNER, an expert-annotated dataset for fine-grained entity recognition in vehicle listings. The dataset includes 659 advertisements from a popular car-se…
- info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
- blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
- info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…