Monolithic vs subjective As pointed out it's hard to gather everyones input to a single result. Rather than have a single fallacy / not fallacy rating have each user be able to express (and own) whether a statement is fallacious. In a usual case it would have the result of "95.4% people think this is a false dictonomy". However there is valuable information on cross-correlating what arguments pass which evaluator. You could have functionality to "ignore all evaluators that think this is a fair argument". People could also profilate themselfs as being quality evaluators. There is a problem/feature where the standard of the evaluator need not be rigour. You could for example have a profilic evaluator for each major political leaning. Or you could aggregate the information by cross-referencing proclaimed political identity ie "65% of self-identified democrats think this argument is fair"
Applicability vs context Being able to target already produced texts means there would be wide applicability. However I am a little concerned on selection effects on what makes it as a "thing to scrutinize". This kind of thing would be effective about small isolated arguments. However politicians that fit their arguments to fit the situation they are presented in could be wrongly presented in being judged outside of that speech situation. Maybe they know that there are better / more valid arguments for their position but choose to utter those they know their audience can relate to. Bringing those arguments under a close scrutiny would be to partly miss the point. I guess part of the idea would be to apply pressure to always use arguments that could pass harsher standards? However I can see many downsides to that. I would rather have all the arguments to be processed to be explicitly (re)created in the context of the website. Then it would be clear that everybody involved respects the clean play attitude and that the arguments are meant to be elaborate and precise. This could mean that only the core and essential points would be covered. That is, it would not be a witch hunt to harass other medias but be an internal matter.
explicitness vs summary score I would have each argument input in a special language/notation that forces every argument to be explicit and computer readable. The arguments would not be prose but collections and networks of semantic tokens. This would provide human language independence ie french and english users would render the tokens in their language but they would be manipulating the same exact ones when one makes a claim in french it would be accessible to the english user too. With the guarantee of computer readableness you could things like compare the axioms of two users and point where contradict, at such a point a discussion is possible. You could then track how often did those discussion shift opinions and which arguments were effective at which populations / belief bases. This could easily be rendered a tool for anti-knowledge seeking testing which manipulations work the best. If such a reduction is not done the meaning of any end result will be a bit nebulous. Its meaning would depend on the process by which it is produced and it would mask approval of a group in the guise of numeric inarguable data. If the vision of what the "clean play" consist off it could be useful but I doubt there is a single axis that would be so critically important to track. I would rather have metrics that tell stuff but don't give a conclusion than reach a conclusion I am not sure what it tells.
The public debate is rife with fallacies, half-lies, evasions of counter-arguments, etc. Many of these are easy to spot for a careful and intelligent reader/viewer - particularly one who is acquainted with the most common logical fallacies and cognitive biases. However, most people arguably often fail to spot them (if they didn't, then these fallacies and half-lies wouldn't be as effective as they are). Blatant lies are often (but not always) recognized as such, but these more subtle forms of argumentative cheating (which I shall use as a catch-all phrase from now on) usually aren't (which is why they are more frequent).
The fact that these forms of argumentative cheating are a) very common and b) usually easy to point out suggests that impartial referees who painstakingly pointed out these errors could do a tremendous amount of good for the standards of the public debate. What I am envisioning is a website like factcheck.org but which would not focus primarily on fact-checking (since, like I said, most politicians are already wary of getting caught out with false statements of fact) but rather on subtler forms of argumentative cheating.
Ideally, the site would go through election debates, influential opinion pieces, etc, more or less line by line, pointing out fallacies, biases, evasions, etc. For the reader who wouldn't want to read all this detailed criticism, the site would also give an overall rating of the level of argumentative cheating (say from 0 to 10) in a particular article, televised debate, etc. Politicians and others could also be given an overall cheating rating, which would be a function of their cheating ratings in individual articles and debates. Like any rating system, this system would serve both to give citizens reliable information of which arguments, which articles, and which people, are to be trusted, and to force politicians and other public figures to argue in a more honest fashion. In other words, it would have both have an information-disseminating function and a socializing function.
How would such a website be set up? An obvious suggestion is to run it as a wiki, where anyone could contribute. Of course, this wiki would have to be very heavily moderated - probably more so than Wikipedia - since people are bound to disagree on whether controversial figures' arguments really are fallacious or not. Presumably you will be forced to banish trolls and political activists on a grand scale, but hopefully this wouldn't be an unsurmountable problem.
I'm thinking that the website should be strongly devoted to neutrality or objectivity, as is Wikipedia. To further this end, it is probably better to give the arguer under evaluation the benefit of the doubt in borderline cases. This would be a way of avoiding endless edit wars and ensure objectivity. Also, it's a way of making the contributors to the site concencrate their efforts on the more outrageous cases of cheating (which there are many of in most political debates and articles, in my view).
The hope is that a website like this would make the public debate transparent to an unprecedented degree. Argumentative cheaters thrive because their arguments aren't properly scrutinized. If light is shone on the public debate, it will become clear who cheats and who doesn't, which will give people strong incentives not to cheat. If people respected the site's neutrality, its objectivity and its integrity, and read what it said, it would in effect become impossible for politicians and others to bullshit the way they do today. This could mark the beginning of the realization of an old dream of philosophers: The End of Bullshit at the hands of systematic criticism. Important names in this venerable tradition include David Hume, Rudolf Carnap and the other logical positivists, and not the least, the guy standing statue outside my room, the "critical rationalist" (an apt name for this enterprise) Karl Popper.
Even though politics is an area where bullshit is perhaps especially common, and one where it does an exceptional degree of harm (e.g. vicious political movements such as Nazism are usually steeped in bullshit) it is also common and harmful in many other areas, such as science, religion, advertising. Ideally critical rationalists should go after bullshit in all areas (as far as possible). My hunch is, though, that it would be a good idea to start off with politics, since it's an area that gets lots of attention and where well-written criticism could have an immediate impact.