Sitemap

A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.

Pages

Posts

Future Blog Post

less than 1 minute read

Published:

This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.

Blog Post number 4

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 3

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 2

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 1

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

portfolio

publications

Extracting an English-Persian Parallel Corpus from Comparable Corpora

A Karimi, E Ansari, BS Bigham

Published in LREC 2018, 2018

We propose a bidirectional method to extract parallel sentences from document-aligned English and Persian Wikipedia.

Recommended citation: Akbar Karimi, Ebrahim Ansari, and Bahram Sadeghi Bigham. 2018. Extracting an English-Persian Parallel Corpus from Comparable Corpora. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
Download Paper

A Novel Region of Interest Extraction Layer for Instance Segmentation

L Rossi, A Karimi, A Prati

Published in ICPR 2020, 2020

We propose GRoIE, a Generic RoI Extractor that uses all FPN layers with non-local blocks and attention, integrating seamlessly into two-stage architectures and improving detection by up to 1.1% AP and instance segmentation by 1.7% AP.

Recommended citation: Leonardo Rossi, Akbar Karimi, and Andrea Prati. 2021. A Novel Region of Interest Extraction Layer for Instance Segmentation. In 2020 25th International Conference on Pattern Recognition (ICPR). IEEE.
Download Paper

Uniparma at SemEval-2021 Task 5: Toxic Spans Detection Using CharacterBERT and Bag-of-Words Model

A Karimi, L Rossi, A Prati

Published in SemEval 2021, 2021

We detect toxic spans by combining CharacterBERT, which handles misspelled toxic words through character-level embeddings, with a bag-of-words method that ensures frequently used toxic words are labeled.

Recommended citation: Akbar Karimi, Leonardo Rossi, and Andrea Prati. 2021. UniParma at SemEval-2021 Task 5: Toxic Spans Detection Using CharacterBERT and Bag-of-Words Model. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 220–224, Online. Association for Computational Linguistics.
Download Paper

Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing

L Rossi, A Karimi, A Prati

Published in CAIP 2021, 2021

We propose R³-CNN, which replaces cascade architectures with a loop mechanism and recursive IoU-based re-sampling, surpassing the HTC model on COCO while significantly reducing the number of parameters.

Recommended citation: Leonardo Rossi, Akbar Karimi, and Andrea Prati. 2021. Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing. In Computer Analysis of Images and Patterns (CAIP 2021).
Download Paper

Improving BERT Performance for Aspect-Based Sentiment Analysis

A Karimi, L Rossi, A Prati

Published in ICNLSP 2021, 2021

We propose two simple modules, Parallel Aggregation and Hierarchical Aggregation, on top of BERT for Aspect Extraction and Aspect Sentiment Classification, improving performance without further training of the BERT model.

Recommended citation: Akbar Karimi, Leonardo Rossi, and Andrea Prati. 2021. Improving BERT Performance for Aspect-Based Sentiment Analysis. In Proceedings of the 4th International Conference on Natural Language and Speech Processing (ICNLSP 2021), pages 39–46, Trento, Italy. Association for Computational Linguistics.
Download Paper

Improving Localization for Semi-Supervised Object Detection

L Rossi, A Karimi, A Prati

Published in ICIAP 2022, 2022

We add a bounding box localization classification task to the Mean Teacher framework to better filter pseudo-labels, and show that box regression on unlabeled data helps as much as classification, improving SSOD on COCO by 1.14% AP.

Recommended citation: Leonardo Rossi, Akbar Karimi, and Andrea Prati. 2022. Improving Localization for Semi-Supervised Object Detection. In Image Analysis and Processing – ICIAP 2022, pages 516–527, Lecce, Italy. Springer.
Download Paper

Aspect-Based Emotion Analysis and Multimodal Coreference: A Case Study of Customer Comments on Adidas Instagram Posts

L De Bruyne, A Karimi, O De Clercq, A Prati, V Hoste

Published in LREC 2022, 2022

We present a multimodal dataset of 4,900 comments on 175 Instagram images annotated for aspect-based emotion analysis, and find that aspect and emotion classification benefit little from multimodal coreference resolution.

Recommended citation: Luna De Bruyne, Akbar Karimi, Orphee De Clercq, Andrea Prati, and Veronique Hoste. 2022. Aspect-Based Emotion Analysis and Multimodal Coreference: A Case Study of Customer Comments on Adidas Instagram Posts. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 574–580, Marseille, France. European Language Resources Association.
Download Paper

Self-Balanced R-CNN for Instance Segmentation

L Rossi, A Karimi, A Prati

Published in JVCIR, 2022

We propose SBR-CNN, an evolution of HTC with loop mechanisms for box and mask refinement and an improved GRoIE, addressing IoU and feature-level imbalances and reaching 45.3% / 41.5% AP on COCO with a ResNet-50 backbone.

Recommended citation: Leonardo Rossi, Akbar Karimi, and Andrea Prati. 2022. Self-Balanced R-CNN for Instance Segmentation. Journal of Visual Communication and Image Representation.
Download Paper

CAISA@SMM4H’22: Robust Cross-Lingual Detection of Disease Mentions on Social Media with Adversarial Methods

A Karimi, L Flek

Published in SMM4H 2022, 2022

We apply adversarial data augmentation in the input and embedding spaces to BioBERT for detecting disease mentions in Spanish tweets, outperforming a vocabulary-based baseline, with augmentation especially helpful in low-data settings.

Recommended citation: Akbar Karimi and Lucie Flek. 2022. CAISA@SMM4H’22: Robust Cross-Lingual Detection of Disease Mentions on Social Media with Adversarial Methods. In Proceedings of the Seventh Workshop on Social Media Mining for Health Applications, Workshop & Shared Task, pages 168–170, Gyeongju, Republic of Korea. Association for Computational Linguistics.
Download Paper

CAISA at SemEval-2023 Task 8: Counterfactual Data Augmentation for Mitigating Class Imbalance in Causal Claim Identification

A Karimi, L Flek

Published in SemEval 2023, 2023

We introduce a counterfactual data augmentation method based on verb replacement for identifying medical claims, yielding significant relative improvement on the minority class compared to three other augmentation techniques.

Recommended citation: Akbar Karimi and Lucie Flek. 2023. CAISA at SemEval-2023 Task 8: Counterfactual Data Augmentation for Mitigating Class Imbalance in Causal Claim Identification. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), pages 2118–2123, Toronto, Canada. Association for Computational Linguistics.
Download Paper

Do Multilingual Large Language Models Mitigate Stereotype Bias?

S Nie, M Fromm, C Welch, R Görge, A Karimi, J Plepi, N Mowmita, N Flores-Herr, M Ali, L Flek

Published in C3NLP Workshop @ ACL, 2024

We investigate the effect of multilingual training on bias mitigation by systematically training six LLMs of identical size (2.6B parameters) and architecture: five monolingual models (English, German, French, Italian, and Spanish) and one multilingual model trained on an equal distribution of data across these languages, all using publicly available data.

Download Paper

Exploring Robustness of Multilingual LLMs on Real-World Noisy Data

A Aliakbarzadeh, L Flek, A Karimi

Published in Eighth Widening NLP Workshop (WiNLP 2024) Phase II, 2024

We study how real-world spelling mistakes from Wikipedia edit history affect 9 multilingual language models across NLI, NER, and intent classification in 6 languages, finding a 2.3 to 4.3 point performance gap, with mT5 models the most robust.

Recommended citation: Amirhossein Aliakbarzadeh, Lucie Flek, and Akbar Karimi. 2024. Exploring Robustness of Multilingual LLMs on Real-World Noisy Data. In Eighth Widening NLP Workshop (WiNLP 2024) Phase II.
Download Paper

Explainable Hallucination through Natural Language Inference Mapping

WF Chen, Z Zhao, A Karimi, L Flek

Published in Findings of ACL, 2025

We introduce HaluMap, a training-free, model-agnostic framework that detects hallucinations by mapping NLI entailment and contradiction relations between inputs and outputs, outperforming other training-free NLI-based methods by five points with interpretable explanations.

Recommended citation: Wei-Fan Chen, Zhixue Zhao, Akbar Karimi, and Lucie Flek. 2025. Explainable Hallucination through Natural Language Inference Mapping. In Findings of the Association for Computational Linguistics: ACL 2025, pages 1888–1896, Vienna, Austria. Association for Computational Linguistics.
Download Paper

Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion

L Fischbach, A Karimi, C Kleen, A Lameli, L Flek

Published in Interspeech 2025, 2025

We use Retrieval-based Voice Conversion to map recordings to a single target speaker, reducing speaker variability so models focus on dialectal features, improving low-resource German dialect classification alone and combined with other augmentations.

Recommended citation: Lea Fischbach, Akbar Karimi, Caroline Kleen, Alfred Lameli, and Lucie Flek. 2025. Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion. In Proc. Interspeech 2025, pages 2780–2784.
Download Paper

EDAudio: Easy Data Augmentation for Dialectal Audio

L Fischbach, A Karimi, A Lameli, L Flek

Published in RANLP 2025, 2025

We evaluate lightweight audio augmentation techniques on recordings from 20 German dialects, finding that frequency-based methods, especially frequency masking, consistently help while time masking or speaker-based insertion can hurt.

Recommended citation: Lea Fischbach, Akbar Karimi, Alfred Lameli, and Lucie Flek. 2025. EDAudio: Easy Data Augmentation for Dialectal Audio. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 363–368, Varna, Bulgaria.
Download Paper

Colliding with Adversaries: A Challenge on Robust Learning in High Energy Physics at ECML PKDD 2025

T Saala, L Flek, A Karimi, PA Jung, A Schmidt, P Soldin, D Stefanopoulos, A Voskou, U Willemsen, C Wiebusch, M Schott

Published in ECML PKDD 2025, 2025

We describe the Colliding with Adversaries challenge, with tasks on generating adversarial examples against a jet-classification model and on building models robust to unseen attacks, using simulated CMS collision data.

Recommended citation: Timo Saala, Lucie Flek, Akbar Karimi, Philipp Alexander Jung, Alexander Schmidt, Philipp Soldin, Dimitris Stefanopoulos, Andreas Voskou, Ulrich Willemsen, Christopher Wiebusch, and Matthias Schott. 2026. Colliding with Adversaries: A Challenge on Robust Learning in High Energy Physics at ECML PKDD 2025. In Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2025), Communications in Computer and Information Science, pages 486–493. Springer.
Download Paper

Enforcing Fundamental Relations via Adversarial Attacks on Input Parameter Correlations

L Flek, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, C Wiebusch

Published in Computing and Software for Big Science, 2025

We present the Random Distribution Shuffle Attack (RDSA), which targets correlations between observables rather than individual features, and show that adversarial training with it improves classification on particle physics and five other tasks.

Recommended citation: Lucie Flek, Philipp Alexander Jung, Akbar Karimi, Timo Saala, Alexander Schmidt, Matthias Schott, Philipp Soldin, and Christopher Wiebusch. 2025. Enforcing Fundamental Relations via Adversarial Attacks on Input Parameter Correlations. Computing and Software for Big Science, 9(1), 19.
Download Paper

Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction

S Rawat, L Flek, A Karimi

Published in WASP Workshop @ IJCNLP-AACL 2025, 2025

We build a multi-task SciBERT system for classifying telescope references, semantic attributes, and instrument mentions in astronomy papers, using stochastic segment sampling and majority voting, which significantly outperforms an open-weight GPT baseline.

Recommended citation: Shivam Rawat, Lucie Flek, and Akbar Karimi. 2025. Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction. In Proceedings of the Third Workshop for Artificial Intelligence for Scientific Publications, pages 195–200, Mumbai, India and virtual. Association for Computational Linguistics.
Download Paper

Shapes are not enough: CONSERVAttack and its use for finding vulnerabilities and uncertainties in machine learning applications

P Bechtle, L Flek, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, C Wiebusch, U Willemsen

Published in arXiv preprint, 2026

We propose CONSERVAttack, an adversarial attack whose perturbations stay within simulation-versus-data uncertainty bounds, evading standard validation checks in high energy physics while fooling the model, and discuss strategies to mitigate such vulnerabilities.

Recommended citation: Philip Bechtle, Lucie Flek, Philipp Alexander Jung, Akbar Karimi, Timo Saala, Alexander Schmidt, Matthias Schott, Philipp Soldin, Christopher Wiebusch, and Ulrich Willemsen. 2026. Shapes are not enough: CONSERVAttack and its use for finding vulnerabilities and uncertainties in machine learning applications. arXiv preprint arXiv:2603.13970.
Download Paper

Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents

MHA Monfared, L Flek, A Karimi

Published in WASSA Workshop @ EACL 2026, 2026

We propose an agentic data augmentation method for Aspect-Based Sentiment Analysis (ABSA) that uses iterative generation and verification to produce high-quality synthetic training examples.

Recommended citation: Mohammad Hossein Akbari Monfared, Lucie Flek, and Akbar Karimi. 2026. Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents. In The Proceedings for the 15th Workshop on Computational Approaches to Subjectivity, Sentiment Social Media Analysis (WASSA 2026), pages 222–234, Rabat, Morocco. Association for Computational Linguistics.
Download Paper

Can LLM Agents Identify Spoken Dialects like a Linguist?

T Bystrich, L Hamm, MH Akhter, L Fischbach, L Flek, A Karimi

Published in DialRes Workshop @ LREC 2026, 2026

We explore whether LLM agents can classify Swiss German dialects from ASR-generated phonetic transcriptions combined with linguistic resources.

Recommended citation: Tobias Bystrich, Lukas Hamm, Maria Hassan Akhter, Lea Fischbach, Lucie Flek, and Akbar Karimi. 2026. Can LLM Agents Identify Spoken Dialects like a Linguist?. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 83–92, Palma de Mallorca. Association for Computational Linguistics.
Download Paper

Speaker Normalization via Voice Conversion Reveals a Human–Machine Dissociation in Dialect Classification

C Kleen, L Fischbach, A Karimi, L Flek, A Lameli

Published in DialRes Workshop @ LREC 2026, 2026

In perception experiments on nine German dialect regions, we show that voice conversion to a single target speaker leaves human dialect recognition unchanged while significantly improving a deep learning model, revealing a divergence between human and machine speech processing.

Recommended citation: Caroline Kleen, Lea Fischbach, Akbar Karimi, Lucie Flek, and Alfred Lameli. 2026. Speaker Normalization via Voice Conversion Reveals a Human–Machine Dissociation in Dialect Classification. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 177–187, Palma de Mallorca. Association for Computational Linguistics.
Download Paper

MiniFool: Physics-Constraint-Aware Minimizer-Based Adversarial Attacks in Deep Neural Networks

L Flek, O Janik, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, M Thiesmeyer, C Wiebusch, U Willemsen

Published in The European Physical Journal C, 2026

We present MiniFool, a physics-inspired adversarial attack that minimizes a χ²-based test statistic combined with a target-score deviation, testing the robustness of neural network classifiers on IceCube, CMS, and MNIST data, including unlabeled experimental data.

Recommended citation: Lucie Flek, Oliver Janik, Philipp Alexander Jung, Akbar Karimi, Timo Saala, Alexander Schmidt, Matthias Schott, Philipp Soldin, Matthias Thiesmeyer, Christopher Wiebusch, and Ulrich Willemsen. 2026. MiniFool: Physics-Constraint-Aware Minimizer-Based Adversarial Attacks in Deep Neural Networks. The European Physical Journal C, 86(6), 641.
Download Paper

More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists

K Alavi, Z Yeltay, L Flek, A Karimi

Published in Findings of ACL 2026, 2026

We evaluate multi-agent sampling-and-voting on adversarially perturbed math questions, finding that more agents improve accuracy but the robustness gap persists.

Recommended citation: Khashayar Alavi, Zhastay Yeltay, Lucie Flek, and Akbar Karimi. 2026. More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists. In Findings of the Association for Computational Linguistics: ACL 2026, pages 43457–43475, San Diego, California, United States. Association for Computational Linguistics.
Download Paper

talks

teaching

Introduction to Natural Language Processing

Undergraduate course, University of Marburg, 2023

Taught the core concepts of NLP, including TF-IDF, word embeddings, RNNs, and Transformers, as well as evaluation methodologies and the applications of NLP methods in various domains such as conversational systems and computational social science.

Dialog Systems

Undergraduate course, University of Bonn, 2024

Taught the main components of a dialog system, e.g., ASR concepts, NLU, dialog manager, dialog state tracking, & NLG, and the tasks involved in each module.

Dialog Systems

Undergraduate course, University of Bonn, 2025

In the lab sessions, I taught the concepts of LLM agents, agentic frameworks such as SmolAgents, LangChain and LlamaIndex, and how to build a personal chatbot using small models, whose quality is improved with them having the capabilities of internet search and tool use.