Monday, 23 December 2019

How to find outliers in document classification with million documents?

I have million documents which belongs to different classes (100 classes). I want to find outlier documents in each class (which doesn't belong to that class but wrongly classified) and filter them. I can do document similarity using cosine similarity by comparing the tokens of each document. I am not able to apply this to filter the wrongly classified documents for a given class. Example: Consider the 3 classes for simplicity with the documents under them.

ClassA  ClassB  ClassC ... 
doc1    doc2    doc3 
doc4    doc5    doc6 
doc7    doc8    doc9 

How can I figure out effectively and efficiently that doc4(and other similar docs) is wrongly classified in ClassA, so that my training data does not contain outliers?



from How to find outliers in document classification with million documents?

No comments:

Post a Comment