Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Curious: is this performance

I wrote a classifier that's 99.96% accurate on comment spam, and 99.99% if configured to present a captcha when it's close.

on a test dataset? Or new, out-of-sample data? I see those percentages and immediately think "overfitting."

Anyway, best of luck! I'm from Florida, so I'll keep you in mind if I hear from anyone needing ML work.



It's live data from a guestbook[0], which was receiving an average of 8200 spam hits a day. There has been a slight decline in spam over the course of the past week or so - from 300/hour to about 250. In addition, there seem to be more coming in that link to unregistered nonsense domains. I've stopped those for the moment by displaying a captcha on all very short posts.

    com.classifyr.scratch> (current-accuracy (log-since (hours 48)))
    0.99969655
[0] http://thecruxshadows.com/guestbook/


Wow! Excellent performance.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: