Skip to content

dennlinger/roberta-cls-consec

Dennlinger/roberta-cls-consec is a machine learning model.

About dennlinger/roberta-cls-consec

This network has been fine-tuned for the task described in the paper Topical Change Detection in Documents via Embeddings of Long Sequences . The weights are based on RoBERTa-base. The training task is to determine whether two text segments (paragraphs) belong to the same topical section or not . This can be utilized to create a topical segmentation of a document by consecutively predicting the "coherence" of two segments . In our training setup, we had entire paragraphs as samples (or up to 512 tokens across two paragraphs), specifically trained on a Terms of Service data set . Note that this might lead to poor performance on "general" topics, such as news articles,
View model source

Explore

FAQ