Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

PHP find relevance

Tags:

php

lucene

sphinx

Say I have a collection of 100,000 articles across 10 different topics. I don't know which articles actually belong to which topic but I have the entire news article (can analyze them for keywords). I would like to group these articles according to their topics. Any idea how I would do that? Any engine (sphinx, lucene) is ok.

like image 236
Patrick Avatar asked Sep 23 '26 05:09

Patrick


2 Answers

In term of machine learning/data mining, we called these kind of problems as the classification problem. The easiest approach is to use past data for future prediction, i.e. statistical oriented: http://en.wikipedia.org/wiki/Statistical_classification, in which you can start by using the Naive Bayes classifier (commonly used in spam detection)

I would suggest you to read this book (Although written for Python): Programming Collective Intelligence (http://www.amazon.com/Programming-Collective-Intelligence-Building-Applications/dp/0596529325), they have a good example.

like image 76
tszming Avatar answered Sep 25 '26 18:09

tszming


Well an apache project providing maschine learning libraries is Mahout. Its features include the possibility of:

[...] Clustering takes e.g. text documents and groups them into groups of topically related documents. Classification learns from exisiting categorized documents what documents of a specific category look like and is able to assign unlabelled documents to the (hopefully) correct category. [...]

You can find Mahout under http://mahout.apache.org/

Although I have never used Mahout, just considered it ;-), it always seemd to require a decent amount of theoretical knowledge. So if you plan to spend some time on the issue, Mahout would probably be a good starting point, especially since its well documented. But don't expect it to be easy ;-)

like image 38
ftiaronsem Avatar answered Sep 25 '26 18:09

ftiaronsem