<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:creator>Proskurnia, Julia</dc:creator>
  <dc:creator>Mavlyutov, Ruslan</dc:creator>
  <dc:creator>Castillo, Carlos</dc:creator>
  <dc:creator>Aberer, Karl</dc:creator>
  <dc:creator>Cudré-Mauroux, Philippe</dc:creator>
  <dc:date>2017</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">Automatically extracting information from social media is challenging given that social  content is often noisy, ambiguous, and inconsistent. However, as many stories break  on social channels first before being picked up by mainstream media, developing  methods to better handle social content is of utmost importance. In this paper, we  propose a robust and effective approach to automatically identify microposts related to  a specific topic defined by a small sample of reference documents. Our framework  extracts clusters of semantically similar microposts that overlap with the reference  documents, by extracting combinations of key features that define those clusters  through frequent pattern mining. This allows us to construct compact and interpretable  representations of the topic, dramatically decreasing the computational burden  compared to classical clustering and k-NN-based machine learning techniques and  producing highly-competitive results even with small training sets (less than 1'000  training objects). Our method is efficient and scales gracefully with large sets of  incoming microposts. We experimentally validate our approach on a large corpus of  over 60M microposts, showing that it significantly outperforms state-of-the-art  techniques.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://folia.unifr.ch/global/documents/307759</dc:identifier>
  <dc:identifier>https://folia.unifr.ch/documents/307759/files/cud_edf.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>info:eu-repo/semantics/altIdentifier/doi/10.1145/3132847.3133016</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>License undefined</dc:rights>
  <dc:source>CIKM '17: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management. - 2017, p. 457–466</dc:source>
  <dc:subject>info:eu-repo/classification/udc/004</dc:subject>
  <dc:title xmlns:ns1="xml" ns1:lang="en">Efficient document filtering using vector space topic expansion and pattern-mining: the case of event detection in microposts</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_5794</dc:type>
</oai_dc:dc>
