Here, the government is pursuing criminal charges for someone who has taken research papers (funded, by and large, through public research grants) to presumably make them available to the public.
In doing so, he wronged various publishers, who have not financially supported any research, who have not financially supported the scientific review of said research, who have financially gained by (likely) not only charging the original author for the submission but also the universities which provided the infrastructure critical to the majority of research.
Spearheading this noble effort to right the wrong and restore law and order is US attorney Carmen Ortiz, who is cited for the wise words: "Stealing is stealing".
Lets be reasonable here: a decently sized server farm could probably keep the entirety of documents hosted on JSTOR in RAM. Today alone, imgur burned through 50TB of traffic; it delivers a petabyte a week. I'm not going to believe a sob story about how distributing 100KiB PDFs to someone running 'wget -r' is DoSing their systems.
But he recklessly damaged a protected computer. They're going to have to replace the fifty cent lock with a new fifty cent lock. That will cost over fifty US cents!
(Yes folks, it's true: pretty much everything is a federal felony if someone doesn't like you.)
Add in beaucratical paperwork, a multitude of companies bidding the lowest price to replace said lock and the salaries of the team of people who are going to be spending the next 2 years replacing it and it really starts to add up.
Well, the article does note that the documents were 'returned'...
>The indictment alleges that Swartz, at the time a fellow at Harvard University, intended to distribute the documents on peer-to-peer networks. That did not happen, however, and all the documents have been returned to JSTOR.
Should I infer from this that they were found by the swat team in a bag labeled 'Swag'?
> Lets be reasonable here: a decently sized server farm could probably keep the entirety of documents hosted on JSTOR in RAM.
One largeish science publisher I worked with had 9-10TB of (pdf) document data. In addition to that there's a search engine, and image versions (of everything) to allow for look inside. Then there's dynamically generated HTML for online view. Electronic publishers do with large amounts of data and it isn't wholly static content. I don't know much about JSTOR but I think it's safe to assume that the implementation is non-trivial.
> I don't know much about JSTOR but I think it's safe to assume that the implementation is non-trivial.
It's non-trivial, and it's also one of the smallest chunks of their budget. As a (bloated, inefficient, high-salary) non-profit, you can read JSTOR's financial filings for yourself:
They spend ~$4m a year on all computer costs. (To put that in perspective, they spend $1.3m a year on 'travel' & 'conferences, conventions, and meetings'.)
I believe Swartz's point is that 10TB of data is not something you can build a self-perpetuating bureaucracy around, anymore, and that this implementation is, in fact, simple enough that it can be made available and maintained for a tiny fraction of the cost of JSTOR.
I'm not really talking about what Swartz's argument is, but you seem to have missed the point that there's more to being a publisher than holding data.
But I don't see the complication on the tech side.
Getting the scanned copies is likely the most complex portion of this and is outside the realm of the site. The rest is just displaying static content and filling a Solr cluster with OCR data for search which is a trivial task with off the shelf OSS tools. Seems like a 2 week MVP project given how clunky the website is (why on earth does it move to the top of the page when I click the 'next' arrow when viewing the article, arg).
However, he did 'break into' an MIT switch closet to run 'keepgrabbing.py' over a 1Gbit/s connection. He wasn't just downloading 100KiB PDFs, either. He downloaded at least two million documents. The indictment isn't clear exactly how many, and it sounds like he downloaded a lot more than 2 million, to boot. Not all of JSTOR's documents are neat 100KiB PDFs, either: a substantial portion are scanned images (1+ MiB PDFs) from old journals. So, we're looking at the TB range of data.
This is not to say that his intentions were ignoble...
So there is at least 2 million scientific documents that publishers are profiting from withholding.
I'm not generally anti-copyright, but I believe the profits publishers make on scientific publishing are unconscionable - not only do they impede progress, but in many cases (eg, medical research) they cost lives.
> “The criminal investigation and today’s indictment of Mr. Swartz has been directed by the United States Attorney’s Office,” said a statement released by JSTOR on July 19. “It was the government’s decision whether to prosecute, not JSTOR’s. As noted previously, our interest was in securing the content. Once this was achieved, we had no interest in this becoming an ongoing legal matter.”
Yes and hopefully it also means that many people take note and speak up against the law and against his harsh treatment.
And this is not just about the law by the way. The fact that governments around the world can be blackmailed or corrupted by a small number of ruthless publishers is a political issue as well. It's also a credibility issue for researchers to some degree, particularly those who have already made a name for themselves and still play along with this.
Academics are completely free to submit their papers to open access journals, the reason they don't do so is because they want the brand value from publishing in prestigious closed-access journals and because submitting to those journals is free.
Building that brand value has risks and cost money to produce (due to editorial, etc costs) hence the publishers need to recoup that value. Open-access journals to recoup that cost either charge the authors or rely upon a subsidies from wealthy benefactors (universities, industrial sponsors).
There are no editorial costs. Editors are not paid. There are no authorial costs. Authors are not paid. There are a few copy-editing costs. Copy-editors get paid a few hundred bucks per article. There are a few administrative costs. The administrators get paid huge amounts. There is also no risk - the journals were set up a long time ago, and are the way they are due to inertia.
If you imagine for-profit academic journal publishers (especially Elsevier) as anything other than vicious monopolists defending their entrenched positions via copyright law and extortion of university libraries, then you have the wrong idea.
> There are a few administrative costs. The administrators get paid huge amounts.
I'll step over the part about these sentences contradicting each other. A colleague of mine worked as a part-time "managing editor" for the top journal in his field, meaning he was paid to coordinate the flow of submissions to the (unpaid) senior editors, managing the responsibilities/schedule of the editor-in-chief, and a host of other work I haven't quizzed him about. He worked hard, and for not much money, which is somewhat reasonable, as he was effectively an executive assistant. The EIC is expected to regularly travel to annual conferences (in one case, because that's when the journal's editorial board meets), so these costs are paid by the journal (to my understanding). He also does a great deal of evangelist work for the journal internationally. While some of these costs may be paid by honorariums (I'm not clear), they must be paid by somebody, and the journal is a likely candidate for a portion of those costs.
My point is that every time we talk about costs of journals, someone mentions either a) administration is a no cost event or b) administration is an overpriced event. Yes, the reviewing and editing is typically performed by volunteers (the costs of which are borne by their respective employers) There is a substantial amount of work that goes on behind the scenes. I've had only the slightest of peeks behind the curtain, and I'm blown away by how much occurs that I wouldn't have initially guessed.
I'm not saying the repository companies (and some publishers) are not money-grubbing parasites; I'm not sure the evidence is in their favor. But little though it's been, my limited view of journal administration is that it's a non-trivial set of tasks costing more than I originally suspected. Claiming otherwise without experience to the contrary is really just FUD.
I was intending to imply that the massive dollar flow to these administrative costs is not proportional to the value provided. The cost for your friend to fly around was not in any way proportional to the price levied on the journal purchasers and users.
I am very interested to hear more about the activities "behind the curtain", because (being a junior academic at a teaching-oriented institution) I haven't had any interaction with a journal publisher that provided actual value beyond providing a website for me to submit my work, and a branded stamp to certify that my work was correct. In every case, I have had to submit a camera-ready final copy that was typeset by me and typo-checked by either me or the unpaid reviewers, and at no point has anyone approached me or anyone else I know asking us to publish our work in journal X as opposed to journal Y. As far as I know, to the degree anything I have ever written has been read at all, it's because I put the PDFs online myself (technically in violation of the copyright agreement, but a behavior that seems to be tacitly accepted in practice). Of course, even after people read the online PDFs (to the degree those papers get read at all), people cite the paper as if they read it in the paper journal or proceedings...
Typically the commissioning and production editors (as opposed to the academic editors) are paid, as are the copyediting staff, the design staff, the sales and marketing teams (brands don't build themselves), etc.
The average academic journal has 25%-30% profit margin (higher than the rest of the publishing industry), but it can take upto 5-7 years for a new STM journal to build enough of a reputation to break even. The profits from the successful journals have to cover the failure of the others (not unlike the startup industry from an investor perspective).
If you look at an open access journal like PLOS Biology the fundamental economics aren't that different. PLOS Biology charges authors $2900/paper, at about 20 papers an issue that means an issue generates just under $60,000 in revenue (excluding revenue from print sales). Overall PLOS runs at a 20% profit rate but it's not significantly different from other mid-tier closed-access journal publishers.
But what do journals do? I know what they did: they were clearinghouses of current scientific knowledge. Now, in the days of the ArXiv, what do they do?
The only answer I can come up with is: help academics gain tenure and promotion. Except for a few flagships (Science, Nature, Cell, etc) most journals are not current, and are not read. Issues are not received with bated breath by people rushing to find out what the cutting edge is. Instead, everyone who cares has gotten the preprints, and everyone else doesn't care.
Arguing about the profitability of journals feels like arguing about the business models of buggy whip manufacturers.
Here, the government is pursuing criminal charges for someone who broke into locked facilities and committed a textbook case of computer fraud to do something that might be morally defensible. If he weren't illegitimately accessing these articles, it'd be a different story.
If you have a tip about an unsolved murder that you want to report anonymously, and you break into someone's house to do so, good on you for reporting the tip, but you still broke into someone's house.
Are you suggesting the case against Swartz is as serious as a burglary in Cambridge, Massachusetts? Or less serious, since he did not damage a lock or deprive the owner of the use of anything? Or is your argument that this is legitimately a high-profile federal case?
In doing so, he wronged various publishers, who have not financially supported any research, who have not financially supported the scientific review of said research, who have financially gained by (likely) not only charging the original author for the submission but also the universities which provided the infrastructure critical to the majority of research.
Spearheading this noble effort to right the wrong and restore law and order is US attorney Carmen Ortiz, who is cited for the wise words: "Stealing is stealing".
Lets be reasonable here: a decently sized server farm could probably keep the entirety of documents hosted on JSTOR in RAM. Today alone, imgur burned through 50TB of traffic; it delivers a petabyte a week. I'm not going to believe a sob story about how distributing 100KiB PDFs to someone running 'wget -r' is DoSing their systems.