Token Classification
Transformers
Safetensors
English
bert
financial NLP
named entity recognition
sequence labeling
structured extraction
hierarchical taxonomy
XBRL
iXBRL
SEC filings
financial-information-extraction
Instructions to use AAU-NLP/Cal-BERT-SL1000 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AAU-NLP/Cal-BERT-SL1000 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="AAU-NLP/Cal-BERT-SL1000")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("AAU-NLP/Cal-BERT-SL1000") model = AutoModelForTokenClassification.from_pretrained("AAU-NLP/Cal-BERT-SL1000", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -45,10 +45,28 @@ This model was introduced in the paper [HiFi-KPI: A Dataset for Hierarchical KPI
|
|
| 45 |
|
| 46 |
### **Citation**
|
| 47 |
```bibtex
|
| 48 |
-
@
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
}
|
| 54 |
```
|
|
|
|
| 45 |
|
| 46 |
### **Citation**
|
| 47 |
```bibtex
|
| 48 |
+
@inproceedings{aavang-etal-2026-hifi,
|
| 49 |
+
title = "{H}i{F}i-{KPI}: A Dataset for Hierarchical {KPI} Extraction from Earnings Filings",
|
| 50 |
+
author = "Aavang, Rasmus T. and
|
| 51 |
+
Rizzi, Giovanni and
|
| 52 |
+
Tjalk-B{\o}ggild, Rasmus and
|
| 53 |
+
Iolov, Alexandre and
|
| 54 |
+
Zhang, Mike and
|
| 55 |
+
Bjerva, Johannes",
|
| 56 |
+
editor = "Piperidis, Stelios and
|
| 57 |
+
Bel, N{\'u}ria and
|
| 58 |
+
van den Heuvel, Henk and
|
| 59 |
+
Ide, Nancy and
|
| 60 |
+
Krek, Simon and
|
| 61 |
+
Toral, Antonio",
|
| 62 |
+
booktitle = "Proceedings of the Fifteenth Language Resources and Evaluation Conference",
|
| 63 |
+
month = may,
|
| 64 |
+
year = "2026",
|
| 65 |
+
address = "Palma de Mallorca, Spain",
|
| 66 |
+
publisher = "ELRA Language Resource Association",
|
| 67 |
+
url = "https://aclanthology.org/2026.lrec-1.30/",
|
| 68 |
+
doi = "10.63317/2nbsp7zzfb3g",
|
| 69 |
+
pages = "441--455",
|
| 70 |
+
abstract = "Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 2.5K-instance subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured extraction. Finally, a qualitative analysis reveals that extraction errors primarily relate to dates. We open-source all code and data at Anonymous."
|
| 71 |
}
|
| 72 |
```
|