lnu.sePublications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 12) Show all publications
Witschard, D., Kucher, K., Jusufi, I. & Kerren, A. (2026). Extending Visually Guided Extraction of Prevalent Topics. In: Proceedings of the 43rd Eurographics Computer Graphics & Visual Computing Conference (CGVC 2026): . Paper presented at CGVC 2026 - the 43rd Eurographics Computer Graphics & Visual Computing Conference, Nottingham, UK, June 11 - 12, 2026. The Eurographics Association
Open this publication in new window or tab >>Extending Visually Guided Extraction of Prevalent Topics
2026 (English)In: Proceedings of the 43rd Eurographics Computer Graphics & Visual Computing Conference (CGVC 2026), The Eurographics Association , 2026Conference paper, Published paper (Refereed)
Abstract [en]

Obtaining overview of large text document data sets remains a challenging and important task across multiple research and application fields. In this paper, we extend our previously proposed prevalence-aware method for topic extraction and successfully apply it to a substantially larger data set than in the previously published examples. This data set has been constructed by scraping the abstract texts of articles from arXiv that contain the keyword ‘LLM’ (denoting large language models), allowing us to obtain insights on the development of this fast moving field of high research interest. We also introduce new functionality within our prototype visual analytics tool, which allows the analyst to combine the results from several different runs and search for common patterns. By extending the maximum size of the input corpus and by providing new strategies for selecting or compiling the best possible result, we position our methodology and tool as a promising candidate for many real-world analysis tasks and scenarios. The main goal is to provide the analyst with the best possible answer to the question “What are the most prevalent topics within this corpus?" and build the trust for the yielded results.

Place, publisher, year, edition, pages
The Eurographics Association, 2026
Keywords
visual analytics, document analysis
National Category
Human Computer Interaction Natural Language Processing Computer Sciences
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-146600 (URN)10.2312/cgvc.20261007 (DOI)
Conference
CGVC 2026 - the 43rd Eurographics Computer Graphics & Visual Computing Conference, Nottingham, UK, June 11 - 12, 2026
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Note

This work was partially supported through the ELLIIT environment for strategic research in Sweden. The work of Ilir Jusufi was supported in part by the Knowledge Foundation, Sweden, through the project ”Rekryteringar 21, Universitetslektor i spelteknik” under Contract 20210077.

Available from: 2026-05-26 Created: 2026-05-26 Last updated: 2026-07-07Bibliographically approved
Witschard, D., Jusufi, I., Kucher, K. & Kerren, A. (2025). Exploring Similarity Patterns in a Large Scientific Corpus. PLOS ONE, 20(4), Article ID e0321114.
Open this publication in new window or tab >>Exploring Similarity Patterns in a Large Scientific Corpus
2025 (English)In: PLOS ONE, E-ISSN 1932-6203, Vol. 20, no 4, article id e0321114Article in journal (Refereed) Published
Abstract [en]

Similarity-based analysis is a common and intuitive tool for exploring large data sets. For instance, grouping data items by their level of similarity, regarding one or several chosen aspects, can reveal patterns and relations from the intrinsic structure of the data and thus provide important insights in the sense-making process. Existing analytical methods (such as clustering and dimensionality reduction) tend to target questions such as "Which objects are similar?"; but since they are not necessarily well-suited to answer questions such as "How does the result change if we change the similarity criteria?" or "How are the items linked together by the similarity relations?" they do not unlock the full potential of similarity-based analysis—and here we see a gap to fill. In this paper, we propose that the concept of similarity could be regarded as both: (1) a relation between items, and (2) a property in its own, with a specific distribution over the data set. Based on this approach, we developed an embedding-based computational pipeline together with a prototype visual analytics tool which allows the user to perform similarity-based exploration of a large set of scientific publications. To demonstrate the potential of our method, we present two different use cases, and we also discuss the strengths and limitations of our approach.

Place, publisher, year, edition, pages
Public Library of Science (PLoS), 2025
Keywords
Visual Text Analytics, Text Mining, Text Embedding, Network Embedding, Similarity Calculations
National Category
Computer Sciences Human Computer Interaction
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-137304 (URN)10.1371/journal.pone.0321114 (DOI)001488705600008 ()2-s2.0-105003254126 (Scopus ID)
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Note

This work was partially supported through the ELLIIT environment for strategic research in Sweden. The work of Ilir Jusufi was supported in part by the Knowledge Foundation, Sweden, through the project ”Rekryteringar 21, Universitetslektor i spelteknik” under Contract 20210077.

Available from: 2025-03-20 Created: 2025-03-20 Last updated: 2025-05-28Bibliographically approved
Witschard, D., Kucher, K., Jusufi, I. & Kerren, A. (2025). Using Similarity Network Analysis to Improve Text Similarity Calculations. Applied Network Science, 10, Article ID 8.
Open this publication in new window or tab >>Using Similarity Network Analysis to Improve Text Similarity Calculations
2025 (English)In: Applied Network Science, E-ISSN 2364-8228, Vol. 10, article id 8Article in journal (Refereed) Published
Abstract [en]

Similarity-based analysis is a powerful and intuitive tool for exploring large data sets, for instance, for revealing patterns by grouping items by similarity or for recommending items based on selected samples. However, similarity is an abstract and subjective property which makes it hard to evaluate by a purely computational approach. Furthermore, there are usually several possible computational models that could be applied to the data, each with its own strengths and weaknesses. With this in mind, we aim to extend the research frontier regarding what impact the choice of a computational model may have on the results. In this paper, we target the scope of embedding-based similarity calculations on text documents and seek to answer the research question: "How can a better understanding of the continuous similarity distribution captured by different models lead to better similarity calculations on document sets?". We propose a new and generic methodology based on similarity network comparison, and based on this approach, we have developed a computational pipeline together with a prototype visual analytics tool that allows the user to easily assess the level of model agreement/disagreement. To demonstrate the potential of our method, as well as showing its application to real world scenarios, we apply it in an experimental setup using three state-of-the-art text embedding models and three different text corpora. In view of the surprisingly low level of model agreement regarding the data, we also discuss strategies for handling model disagreement.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Embeddings, Text Similarity Calculations, Similarity Networks, Visual Analytics
National Category
Computer Sciences
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-137305 (URN)10.1007/s41109-025-00699-7 (DOI)001467943200001 ()2-s2.0-105000480934 (Scopus ID)
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Note

This work was partially supported through the ELLIIT environment for strategic research in Sweden. The work of Ilir Jusufi was supported in part by the Knowledge Foundation, Sweden, through the project ”Rekryteringar 21, Universitetslektor i spelteknik” under Contract 20210077.

Available from: 2025-03-20 Created: 2025-03-20 Last updated: 2025-05-28Bibliographically approved
Witschard, D., Jusufi, I., Kucher, K. & Kerren, A. (2025). Visually Guided Extraction of Prevalent Topics. Information Visualization, 42(2), 179-198
Open this publication in new window or tab >>Visually Guided Extraction of Prevalent Topics
2025 (English)In: Information Visualization, ISSN 1473-8716, E-ISSN 1473-8724, Vol. 42, no 2, p. 179-198Article in journal (Refereed) Published
Abstract [en]

The sensemaking process of large sets of text documents is highly challenging for tasks such as obtaining a comprehensive overview or keeping up with the most important trends and topics. Even though several established methods for condensation and summarization of large text corpora exist, many of them lack the ability to account for difference in prevalence between identified topics, which in turn impedes quantitative analysis. In this paper, we therefore propose a novel prevalence-aware method for topic extraction, and show how it can be used to obtain important insights from two text corpora with very different content. We also implemented a prototype visual analytics tool which guides the user in the search for relevant insights and promotes trust in the yielded results. We have verified our application by a user study, as well as by a validation run on a data set with previously known topic structure. The results clearly show that our approach is suitable for text mining, that is can be used by non-experts, and that it offers features which makes it an interesting candidate for use in several different analyze scenarios.

Place, publisher, year, edition, pages
Sage Publications, 2025
Keywords
Visual Analytics, Text Mining, Text Embedding, Topic Modelling, Similarity Calculations
National Category
Computer Sciences Human Computer Interaction
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-136101 (URN)10.1177/14738716241312400 (DOI)001408697200001 ()
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Note

This work was partially supported through the ELLIIT environment for strategic research in Sweden. The work of Ilir Jusufi was supported in part by the Knowledge Foundation, Sweden, through the project ”Rekryteringar 21, Universitetslektor i spelteknik” under Contract 20210077.

Available from: 2025-02-09 Created: 2025-02-09 Last updated: 2026-08-31Bibliographically approved
Huang, Z., Witschard, D., Kucher, K. & Kerren, A. (2023). VA + Embeddings STAR: A State-of-the-Art Report on the Use of Embeddings in Visual Analytics. Paper presented at 25th EG Conference on Visualization (EuroVis '23), STAR track, 12-16 June 2023, Leipzig, Germany. Computer graphics forum (Print), 42(3), 539-571
Open this publication in new window or tab >>VA + Embeddings STAR: A State-of-the-Art Report on the Use of Embeddings in Visual Analytics
2023 (English)In: Computer graphics forum (Print), ISSN 0167-7055, E-ISSN 1467-8659, Vol. 42, no 3, p. 539-571Article in journal (Refereed) Published
Abstract [en]

Over the past years, an increasing number of publications in information visualization, especially within the field of visual analytics, have mentioned the term “embedding” when describing the computational approach. Within this context, embeddings are usually (relatively) low-dimensional, distributed representations of various data types (such as texts or graphs), and since they have proven to be extremely useful for a variety of data analysis tasks across various disciplines and fields, they have become widely used. Existing visualization approaches aim to either support exploration and interpretation of the embedding space through visual representation and interaction, or aim to use embeddings as part of the computational pipeline for addressing downstream analytical tasks. To the best of our knowledge, this is the first survey that takes a detailed look at embedding methods through the lens of visual analytics, and the purpose of our survey article is to provide a systematic overview of the state of the art within the emerging field of embedding visualization. We design a categorization scheme for our approach, analyze the current research frontier based on peer-reviewed publications, and discuss existing trends, challenges, and potential research directions for using embeddings in the context of visual analytics. Furthermore, we provide an interactive survey browser for the collected and categorized survey data, which currently includes 122 entries that appeared between 2007 and 2023.

Place, publisher, year, edition, pages
John Wiley & Sons, 2023
Keywords
embedding techniques, distributed representations, visual analytics, visualization
National Category
Computer Sciences
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-120749 (URN)10.1111/cgf.14859 (DOI)001020716600041 ()2-s2.0-85163625612 (Scopus ID)
Conference
25th EG Conference on Visualization (EuroVis '23), STAR track, 12-16 June 2023, Leipzig, Germany
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile CommunicationsWallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2023-05-16 Created: 2023-05-16 Last updated: 2025-05-28Bibliographically approved
Witschard, D., Jusufi, I., Kucher, K. & Kerren, A. (2023). Visually Guided Network Reconstruction Using Multiple Embeddings. In: Proceedings of the 16th IEEE Pacific Visualization Symposium (PacificVis '23), visualization notes track, IEEE, 2023: . Paper presented at 16th IEEE Pacific Visualization Symposium (PacificVis '23), Seoul, Korea, April 18-21, 2023 (pp. 212-216). IEEE
Open this publication in new window or tab >>Visually Guided Network Reconstruction Using Multiple Embeddings
2023 (English)In: Proceedings of the 16th IEEE Pacific Visualization Symposium (PacificVis '23), visualization notes track, IEEE, 2023, IEEE, 2023, p. 212-216Conference paper, Published paper (Refereed)
Abstract [en]

Embeddings are powerful tools for transforming complex and unstructured data into numeric formats suitable for computational analysis tasks. In this paper, we extend our previous work on using multiple embeddings for text similarity calculations to the field of networks. The embedding ensemble approach improves network reconstruction performance compared to single-embedding strategies. Our visual analytics methodology is successful in handling both text and network data, which demonstrates its generalizability beyond its originally presented scope.

Place, publisher, year, edition, pages
IEEE, 2023
Keywords
Graph embedding, network embedding, similarity calculations, visual analytics, visualization
National Category
Computer Sciences Human Computer Interaction
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-119859 (URN)10.1109/PacificVis56936.2023.00031 (DOI)2-s2.0-85163367392 (Scopus ID)9798350321241 (ISBN)9798350321258 (ISBN)
Conference
16th IEEE Pacific Visualization Symposium (PacificVis '23), Seoul, Korea, April 18-21, 2023
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Available from: 2023-03-19 Created: 2023-03-19 Last updated: 2025-05-28Bibliographically approved
Witschard, D., Jusufi, I., Martins, R. M., Kucher, K. & Kerren, A. (2022). Interactive Optimization of Embedding-based Text Similarity Calculations. Information Visualization, 21(4), 335-353
Open this publication in new window or tab >>Interactive Optimization of Embedding-based Text Similarity Calculations
Show others...
2022 (English)In: Information Visualization, ISSN 1473-8716, E-ISSN 1473-8724, Vol. 21, no 4, p. 335-353Article in journal (Refereed) Published
Abstract [en]

Comparing text documents is an essential task for a variety of applications within diverse research fields, and several different methods have been developed for this. However, calculating text similarity is an ambiguous and context-dependent task, so many open challenges still exist. In this paper, we present a novel method for text similarity calculations based on the combination of embedding technology and ensemble methods. By using several embeddings, instead of only one, we show that it is possible to achieve higher quality, which in turn is a key factor for developing high-performing applications for text similarity exploitation. We also provide a prototype visual analytics tool which helps the analyst to find optimal performing ensembles and gain insights to the inner workings of the similarity calculations. Furthermore, we discuss the generalizability of our key ideas to fields beyond the scope of text analysis.

Place, publisher, year, edition, pages
Sage Publications, 2022
Keywords
Text embedding, ensemble methods, text similarity, similarity calculations, visual analytics
National Category
Computer Sciences Natural Language Processing
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-115658 (URN)10.1177/14738716221114372 (DOI)000835467000001 ()2-s2.0-85136447359 (Scopus ID)
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Available from: 2022-08-04 Created: 2022-08-04 Last updated: 2026-04-16Bibliographically approved
Witschard, D. (2022). Towards Multiple Embeddings for Multivariate Network Analysis. (Licentiate dissertation). Linnaeus University Press
Open this publication in new window or tab >>Towards Multiple Embeddings for Multivariate Network Analysis
2022 (English)Licentiate thesis, monograph (Other academic)
Abstract [en]

The study of multivariate networks (MVNs, i.e., large data sets where datapoints have relations to other data points and both these relations and the pointsthemselves can have attributed data) is an important task in many different fields,such as social networks for the humanities, citation networks for bibliometricsand biochemical networks for life sciences. Furthermore, when dealing withvisualization and analysis of MVNs, many open challenges still exist regardingboth computational aspects (i.e., the challenge of computing different metricsof a large-scale MVN) and visual aspects (i.e. the challenge of displaying allthe information of a large-scale MVN in a way that is comprehensible to theuser). In the search for efficient and scalable visual analytics methods, especiallyfor exploratory data analysis, this thesis explores a novel approach of aspectdrivenMVN embedding and the use of ensembles of embeddings for multi-levelsimilarity calculations. Starting from the observation that there already existseveral different embedding techniques for datatypes that are common for realworldMVNs, the main question that we will try to answer is: “Could the useof multiple embeddings provide for new and better solutions for visual analytics onmultivariate networks?" This main question then inspires the formulation of fourmore specific research goals regarding: (1) methods for combining embeddings,(2) the development of a general methodology framework, (3) new visualizationmethods, and (4) proof-of-concept applications for real-world scenarios.The focus of our work lies on similarity-based analysis within the domainsof bibliometrics and scientometrics, and our first major step is to developa methodology for combining several different embeddings (for the sameunderlying data) to augment the quality of similarity calculations. This stepincludes an adaptation of some of the key ideas from ensemble methods to thefield of embeddings, and also an interactive optimization process for finding thebest performing ensembles. Upon this foundation, we develop an aspect-drivenapproach which seeks to divide an underlying MVN into separately embeddableaspects, which in turn allows for the resulting embedding vectors to be used inflexible analysis scenarios with high level of interaction. We then proceed toshow how the concept of similarity-based analysis can be used to obtain valuableinsights to, and a better understanding of, a large set of scientific publications.For this, we introduce the abstract concept of similarity patterns which we use toexpress how a specific set of similarity criteria are distributed over a data set.Furthermore, we present proof-of-concept applications which are designed toallow the user to exploit these similarity patterns at different levels of detail. Wealso show that our proposed methodology is generalizable beyond the scope ofMVNs, and therefore could be applied to other fields as well.

Place, publisher, year, edition, pages
Linnaeus University Press, 2022. p. 72
Series
Lnu Licentiate ; 37
National Category
Computer Systems
Research subject
Computer and Information Sciences Computer Science, Computer Science
Identifiers
urn:nbn:se:lnu:diva-114998 (URN)9789189460980 (ISBN)9789189460997 (ISBN)
Supervisors
Available from: 2022-06-29 Created: 2022-06-29 Last updated: 2023-08-21Bibliographically approved
Witschard, D., Jusufi, I., Martins, R. M. & Kerren, A. (2021). A Statement Report on the Use of Multiple Embeddings for Visual Analytics of Multivariate Networks. In: Christophe Hurter, Helen Purchase, José Braz, and Kadi Bouatouch (Ed.), Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2021) - Volume 3: IVAPP: . Paper presented at International Conference on Information Visualization Theory and Applications (IVAPP), Virtual Conference, 8-10 February, 2021 (pp. 219-223). SciTePress, 3
Open this publication in new window or tab >>A Statement Report on the Use of Multiple Embeddings for Visual Analytics of Multivariate Networks
2021 (English)In: Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2021) - Volume 3: IVAPP / [ed] Christophe Hurter, Helen Purchase, José Braz, and Kadi Bouatouch, SciTePress, 2021, Vol. 3, p. 219-223Conference paper, Published paper (Refereed)
Abstract [en]

The visualization of large multivariate networks (MVN) continues to be a great challenge and will probably remain so for a foreseeable future. The field of Multivariate Network Embedding seeks to meet this challenge by providing MVN-specific embedding technologies that targets different properties such as network topology or attribute values for nodes or links. Although many steps forward have been taken, the goal of efficiently embedding all aspects of a MVN remains distant. This position paper contrasts the current trend of finding new ways of jointly embedding several properties with the alternative strategy of instead using, and combining, already existing state-of-the-art single scope embedding technologies. From this comparison, we argue that the latter strategy provides a more generic and flexible approach with several advantages. Hence, we hope to convince the visual analytics community to invest more work in resolving some of the key issues that would make this methodology possible.

Place, publisher, year, edition, pages
SciTePress, 2021
Keywords
Multivariate Network, Visualization, Visual Analytics, Embedding, Methodology
National Category
Computer Sciences Computer and Information Sciences
Research subject
Computer Science, Information and software visualization; Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-100121 (URN)10.5220/0010314602190223 (DOI)000661282300021 ()2-s2.0-85102971461 (Scopus ID)9789897584886 (ISBN)
Conference
International Conference on Information Visualization Theory and Applications (IVAPP), Virtual Conference, 8-10 February, 2021
Available from: 2021-01-17 Created: 2021-01-17 Last updated: 2026-04-16Bibliographically approved
Witschard, D., Jusufi, I. & Kerren, A. (2021). Dynamic Ranking of IEEE VIS Author Importance. In: Poster Abstract, IEEE Visualization and Visual Analytics (VIS '21): . Paper presented at IEEE VIS: Visualization & Visual Analytics, Virtual. IEEE
Open this publication in new window or tab >>Dynamic Ranking of IEEE VIS Author Importance
2021 (English)In: Poster Abstract, IEEE Visualization and Visual Analytics (VIS '21), IEEE, 2021Conference paper, Poster (with or without abstract) (Refereed)
Abstract [en]

The ranking of authors is an important task within the field of sci- entometrics, and several different methods and criteria exist. In this poster abstract, we present an interactive visualization approach for exploring combinations of several different ranking criteria for a given set of publications and its associated co-author network. Ourvisualization tool allows the user to gain insights into the relative importance of individual authors as well as into the interdependency of different ranking criteria.

Place, publisher, year, edition, pages
IEEE, 2021
National Category
Computer Sciences Human Computer Interaction
Research subject
Computer Science, Information and software visualization
Identifiers
urn:nbn:se:lnu:diva-110952 (URN)
Conference
IEEE VIS: Visualization & Visual Analytics, Virtual
Available from: 2022-03-24 Created: 2022-03-24 Last updated: 2023-08-21Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-6150-0787

Search in DiVA

Show all publications