Any hints as to how to prevent site abuse (by I think, mathematical models and statistics) can be posted hereemmem wrote:Maybe there is another kind of ranking users. Depending of how much their data can be trusted.
A Woombelist : Having an amount of hits out of proportion
A Serialist : Having a certain percentage (for instance more than 50% of the bills) with more or less consecutive serials.
A Heatofeuroper : someone who enters bills in several counties and where it’s obvious he couldn’t have visited them all
An Errorist : a user who has exceeded a certain percentage of errors between serial and printer codes.
I do not know if there is a model for bills, but there is a model for coins ...
If you know how bills and coins should spread, you can detect all data manipulations, and mark them as suspicious ...
The model I refer to can be found at
If we assume the model is correct and can be applied to bills, we can calculate the odds for a real hit (at least theoretically ... check out the site : I think there is software available to do this for you)Skylimit wrote: There is a theory about the distribution of euro-coins, which uses some heavy maths to explain how coins mix, and how long it takes
http://www.wiskgenoot.nl/eurodiffusie/theorie/eurod.pdf
The model should be valid to a certain extent for bills to ... remark : the 2 euro coins seem to travel faster than the little ones ... so one may assume that 5 euro bills go even faster ... and what about the 500 euro bills ... The report was written in april, and is based on a website, where only the distribution mix was measured
If we know what we expect, we also know what we don't expect ...
Woombelism : an unusual amount of hits can be quantified
Serialism : is a different problem. a serialm check can be run for specific users which enter suspicious data. The check could search for consecutive numbers, or a pattern in the numbers, which is impossible if we expect a complete randomized input. More advanced cheaters will then probably start randomizing their input, making it impossible to do the check... but over time, their hits will not match the pattern we expect ...
To name one thing : a user claiming to generate his input in Germany must be aware that his bills distribution mix, must match the mix that is specific to Germany in that specific month ... Hey, don't we just have this information at our fingertips ... wasn't that the original purpose of the site ... Randomized numbers will therefore also produce a pattern we do not expect.
Heat of Europe problem : This user must be aware that his input, he claims to generate in countries all over the world must match the distribution mix that is specific to that country (provided enough data exists ...)
Intercontinentals : Users claiming to be from another continent but really are not, should be aware of the timezone ... A user claiming to be from the Westcoast, should be aware not to input his data while average people are sleeping...
Errorist-problem : This site could collect the probability that a user generates an error. We then know what to expect, so we also know what not to expect. Users that generate significantly more errors than the average, can be spotted electronically
There is one problem however : in a normal distribution, there are always extremes. Banning the extremes from the database also equals manipulating the data. The very data we use to find out what pattern we expect ...
Maybe the study group http://www.wiskgenoot.nl/eurodiffusie has some ideas for solving your problem ...
Provided we accept that extremes cannot be banned from the database, because this also corrupts the results, we can still implement the skylimit-tree-concept, which is legendary by now and solves all problems. You can't fake a network ...
Kind regards



