Comments (6)
Thank you for the kind words!
Just to make sure I understand correctly, do you mean that you want the frequencies of the topics in the dataset? If so, you could get that with .get_topic_info()
.
from bertopic.
Dear Maarten,
Thanks for your quick reply!
Emm, what we are looking for is a value somehow can tell how important a topic contributes to the dataset.
Thank you.
Li
from bertopic.
Emm, what we are looking for is a value somehow can tell how important a topic contributes to the dataset.
Just to make sure that I understand correctly, how would you define "important"? To me, it seems like the most common topic would be the one that contributes most to the dataset.
from bertopic.
Dear Maarten,
Thanks again for your response. And I apologize for my vague question. Yes, I agree with you. The most common topic would be the one that contributes most to a dataset. So, I can pull the information by calling .get_topic_info(), am I right? Thank you again.
Li
from bertopic.
Yes, you can extract the frequency with .get_topic_info()
.
from bertopic.
Thank you so much Maarten!
from bertopic.
Related Issues (20)
- Supervised topic model generating different topics to training data HOT 3
- Where is the full data set of embeddings? HOT 3
- Visualization in html page HOT 1
- Guided Modeling: Problem with seed_topic_list HOT 2
- Utilizing the GPU of MacBook Pro M3 to accelerate the process of fit_transform HOT 1
- Can't reproduce same results when using cuml version of UMAP and HDBSCAN HOT 3
- approximate_distribution returns only 0s HOT 5
- Feature (Watsonx): representations using Llama-3-70b and Mixtral-8x7b HOT 1
- Which hyper parameter mostly influence the number of topics for Chinese texts? HOT 3
- Zero-Shot Topic Modelling and Topics Over Time HOT 1
- Loading of saved model returns Error: "This BERTopic instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator."
- Creating representations using IBM Watsonx LLMs HOT 5
- c_tf_idf_ is None when using zero shot topic modeling. HOT 1
- Issue with Scikit-learn 1.5.0
- Error at Combining clustered topics with the zeroshot model HOT 2
- Compare LDA, NMF, LSA with BERTopic (w/ embedding: all-MiniLM-L6-v2 + dim_red: UMAP + cluster: HDBSCAN) HOT 1
- AttributeError: 'BertModel' object has no attribute 'attn_implementation' #30965 HOT 3
- Zeroshot Topic Modeling With no Embedding Model HOT 1
- Extending ".visulize_document_datamap" with "label_over_points"-flag
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from bertopic.