It will sound odd, but I have recently had a great time editing URLs. Robert Shoemaker and I have have just finished a book for CUP, derived from the London Lives project, and called - London Lives: Poverty, Crime and the Making of a Modern City, 1690-1800. It is a long book (170,000 words) and each quote and reference in it is linked via a URL to the original document or article, book or web-resource used as evidence or to contextualize the argument. It will be published as both an ebook and in hard copy, and the links need to be robust, and secure. My estimate is that there are in the region of 4,000 URLs included in the manuscript (which was written collaboratively in PMWiki). In the end, I found that I could identify an appropriate link for 98% of all footnote references, but then had to eliminate around 10% of these, as the relevant URL was just not useable. The book took some nine years, and I am glad it is finished.
One of my final jobs was editing those 4000 URLs. It took about three months work, spread over the last year, and I have just finished spending a week or so confirming what I hope will be their final form. When I have told people about this work many have looked incredulous and suggested that this is the sort of technical implementation process that should be left to others. A couple of otherwise nice people have suggested I dump this job on the shoulders of the nearest PhD student. But for myself, it is precisely the kind of thing that an author should do for themselves. And in doing it, two things kept coming to mind. First was how the role of the scholar in creating a rigorous academic apparatus is a central part of the intellectual journey that academic writing involves - and that we should see the implementation of the online version of this in the light of the precise writing of footnotes and references that mark out good scholarship. And second, that URLs encode a system of design and intent, online architecture and system of access, that signal the quality and permanence (the academic credibility and perceived audience) of historical materials online. And that just as we have always sorted and judged scholarship by its form, we should think a bit harder about how the form of a URL can let us interrogate online materials.
On the first point, I do not know of much discussion of the joys of this kind of academic slog. There is a lot of good writing on research and archives (by Carolyn Steedman and Arlette Farge among many others), on writing and thinking, but no-one talks much about the painstaking labour that goes in to turning a rough draft in to a final finished piece of scholarship. And here I am really talking about generating accurate and fully comprehensive footnotes that reflect both the material cited, and the research journey that resulted in the main text. This has become much easier with online catalogs and citation management packages, but nevertheless remains laborious and a reflection of our collective and individual commitment to a particular kind of evidenced discussion. But for me it also represents my favourite compromise. The writing of history is a wonderfully imaginative and creative process. And in some respects we wish to judge the product of history writing as art. Is it enjoyable to read? Is it convincing? Does it do the job of good writing in liberating the readers' imagination? In making these judgements we tend to appeal to a notion of 'value' that is cultural and that privileges dominant forms of authority. This aspect of judgement is essentially romantic; with all the implications for western and elite hegemony embedded in that idea. At the same time history writing is the result of simple hard work of a more technical kind - in the archives, in collating and collecting, re-ordering and interrogating data. And it is valuable because it encompasses that hard work. The beauty of the academic apparatus is that it evidences this and in the process generates a different measure of value. In other words it is where quality is tied to a 'labour theory of value'. I love the academic slog because it is where un-moored judgement is tied down to hard labour; and where value can be universalized in a common human experience (work). In other words I really enjoyed editing 4000 URLs precisely because in them and their associated footnotes lies a claim to and evidence of the hard labour that underpins the book itself.
At the same time, the process also taught me to read URLs differently. Clearly coders and web designers do this as a matter of course. But I am a historian and want to read URLs as a scholar, rather than as a programmer or designer. And for me, the important thing is that URLs embed the structure of a site, making it plain to see for anyone willing to look hard; and that they are made up of both the character of a library reference, and a command directed at the new technology of discovery - the Internet . There are just lots of different types of URL.
There are 'Search URLs' that include all the elements that take the user past a collection to a specific object, but don't let you go directly there without the query. And there are URLs that encode a cataloging hierarchy. There are URLs that sift data, or work in your browser to change the data delivered, highlighting phrases or sifting material. And there are URLs that encode licensing, passwords, and access information. It is easy enough to find that the whole search journey that took you from a library catalog to an individual item is encoded directly in the URL, and even personalized to you, the machine you are using, or the forms of access you can deploy. It is easy to find URLs that run on for hundreds of characters, each element divided by a '&' or a '%', or such.
But in creating robust reproducible links to credible historical materials most of these URLs are at least problematic if not useless. If they include details for institutional access, or session information, they cannot be re-used by someone else. These URLs are friable and fragile things and not fit for scholarly purposes. And as a result, for the London Lives book we have been forced to eliminate all the links we originally hoped to include to forty or fifty different sites. To take a single example, most archives structure their online collections with search in mind, making it difficult to link to a single item. I spent a lot of time finding the catalog entry for every manuscript we cited in the London Metropolitan Archives, and Westminster Archives Centre, only to regretfully strip out the links when confronted by a complex URL that just did not look credible as a long term citation of the item itself.
Even in its simplest, and in the form recommended by the site for sharing a link, a London Metropolitan Archives URL looks like this:
http://search.lma.gov.uk/scripts/mwimain.dll/144/LMA_OPAC/web_detail/REFD+P69~2FBRI~2FB~2F001~2FMS06554~2F004?SESSIONSEARCH
Since we had consulted these items in their physical form in any case, it did not seem too problematic to leave out these links, but a shame nevertheless. And likewise, with paywall material there seemed little point in dangling real access, and the promise of credible evidence, before the eyes of readers who would not be able to go beyond the login screen. It seemed better to cite a specific item in combination with a general (unlinked) URL and date of consultation as reflecting our own research journey, rather than to promise access when we could not deliver it.
With few exceptions the URLs that have been retained (and there are still 4000 of them) address specific items with a specific ID, and usually run to 20 to 40 characters. DOIs are not bad once you figure out their structure and reformulate them as they should be, rather than the way they are normally cited on journal web pages.
dx.doi.org/10.1353/sec.2010.0268
And Google Books creates a very nice URL once you strip out all the complex formatting instructions that are normally generated as part of a search and inserted after the main ID. This is what a Google Books' URL looks like if you were to use the 'search' version:
http://books.google.co.uk/books?id=1sMJGt7_rTAC&printsec=frontcover&dq=%22Prosecution+and+Punishment:+Petty+Crime+and+the+Law%22&hl=en&sa=X&ei=rrzGUq_aDsSy7Aa_9YGQCg&redir_esc=y#v=onepage&q=%22Prosecution%20and%20Punishment%3A%20Petty%20Crime%20and%20the%20Law%22&f=false
And this URL will take to the same book:
books.google.co.uk/books?id=1sMJGt7_rTAC
And the Eighteenth-century Short Title Catalog generates some of the most elegant URLs I have found:
estc.bl.uk/T174945
And to a lesser extent, so does the Ethos collection of doctoral theses at the British Library.
ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.354762
And London Lives and the Old Bailey Online do pretty well on this score:
www.londonlives.org/browse.jsp?div=LMSMPS501980014
http://www.oldbaileyonline.org/browse.jsp?ref=t17910413-19
In part, I suspect that these issues would all disappear if I had a better sense of the layer of structure that lies beneath the WWW. But for the moment I am keen to have a short, human-readable URL that looks like it will last longer than the session I am currently logged on for. All of which simply takes me back to the joy of academic slogging and the importance of the academic apparatus as something that evidences hard work and opens up scholarship to credible criticism that goes beyond simple romantic appreciation and prejudice.
I know all too well that one of the skills of an academic is the ability to judge a book by its cover and the form of the text it contains. For the online we need to embed URLs into precisely this process - and the joy of all that editing was that at the end of it, I feel I have learned to do just that.
This blog is a space for me to rant in that most seventeenth-century sense of the word; and to cut and paste the ideas and comments that don't seem to fit in more traditional forms of academic publication.
Showing posts with label book history. Show all posts
Showing posts with label book history. Show all posts
Friday, 3 January 2014
Monday, 4 March 2013
OA in the UK
A recent one-day colloquium sponsored by the Institute of Historical Research and the Royal Historical Society was called precisely to bring together major institutional players (Scholarly Societies, Journals, and Publishers) for a conversation about the best ways forward (the Tweet stream is here). The general feeling seems to be that while every well-meaning historian is keen to promote Open Access (a show of hands at the conference confirmed this), the Gold Route, whereby authors and institutions are asked to shoulder the cost of peer review and publishing, is just not workable in the humanities.
There is also the beginnings of what many feel is an apparent solution to the problem. Both the past and present presidents of the Royal Historical Society, and some 21 editors of major humanities journals, signed a letter proposing the imposition of an increased embargo period on the articles in their journals - essentially suggesting they be allowed to have three years in which to make money on their publications before being forced to make them available through Open Access. They also proposed maintaining a two-track system to ensure that overseas and non-academic authors are excluded from the government led requirement for Open Access.
To me this feels like Saint Augustine's plaint: "Grant
me chastity and continence, but not yet." (Confessions 8:17).
Let me be clear, though.
I understand completely the anxieties motivating these institutions and
commentators. A narrowly defined Gold
Route process of the sort privileged by the Finch Report is not workable in the
humanities. The 'author pays' model is
predicated on the direct funding of research by government, and on the
assumption that the consumers of research outputs are the same as the
producers. In the case of history this
is not true.
The vast majority of
historical research and publication is not funded by project grants; and while
a higher proportion is funded through the Universities, and through QR, there
is still a large body of excellent work that is undertaken by independent
scholars, or as part of a self-funded PhD, or by staff in institutions which do
not receive QR funding or participate in the REF. And similarly, all historians seek to reach a
wider audience than most scientists, and imagine their work in the light of a successful 'trade
monograph'; which itself forms a recognised academic achievement.
In other words, I largely agree with the diagnosis that the main
thrust of the Finch Report is unworkable.
Though, of course, the Report does not restrict academics to the single
route to Open Access, and makes it clear that other types of OA (Green Route)
are entirely consistent with the objective of making publicly funded research
available to the public. Following on
from this, I also believe the RCUK policy to cover the new costs entailed through
grants is also largely unworkable, and if poorly executed in pursuit of a
narrow Gold Route form of publication will create issues of fair access, with institutional
meddling in academic decision making and serious problems for post-graduates
and early career academics.
What is missing in all this is any positive model of Open
Access publishing that takes seriously the fundamental interests and values of history
as a discipline, as opposed to the interests of the collection of institutions
and journals that purport to speak for it.
For myself I have a clear sense of what I would like Open Access in
history (and more broadly in the academy) to look like in ten years' time; and
it would have the following characteristics:
- It would be built on the deposit of articles and research data (including notes) in institutional repositories, linked to APIs that allow their content to be re-'published', mashed up and re-used (with acknowledgment).
- The route to 'publication' would include the initial deposit of research materials, followed by the posting of a 'rough draft' for comment and revision, leading to a post-publication peer review system. The author would then be allowed to specify at what revision the 'article' is complete, with perhaps a six month norm for revisions. See for instance the History Working Papers Project.
- Metrics for downloads, re-use and citations for the now online article will be used to generate a measure of scholarly importance. These metrics can include the kind of complex systems for assessing 'authority' (i.e. whose post peer review assessments are worth most) implicit in the Altmetrics movement.
- 'Journals' will be made up of adopted 'articles' that fit their theme and which are moulded to a particular house style in the open peer review process. In this ecology of scholarship journals will take on a new intellectual role in shaping debate and argument, and in defining academic communities, and will have a 'promoting' role rather than a 'publishing' role.
- Academic monographs will be seen as a simple extension of article publication - i.e. either in the form of long articles, or perhaps as a collection of pieces created as 'articles'.
- Genuine 'Trade History' will continue to be sold in the generic forms of biography and narrative etc., but the underlying academic historical content will be available in institutional repositories, while the revisions and adaptions for a popular audience are dealt with separately.
- The co-archiving of secondary writing and notes and research materials will allow for the creation of an increasingly vertically integrated form of writing, in which source material and commentary are connected.
- The costs of maintaining, curating and archiving the system will be borne by the Universities, with savings from the journal and book purchasing costs. A separate tNA or British Library repository will support and archive the work of otherwise unaffiliated scholars.
Getting to this point is not straightforward. But the universities are fully empowered to
use the RCUK funding to beef up their repositories, rather than paying journal
fees. The repositories could also take a first step
towards a more rigorous ecology of scholarship by archiving and making
available the underlying research data we all endlessly collect (and jealously
guard as the capital of an ego driven system of professional advancement).
At the moment most public debate and effort seems to be
devoted to preserving the current business model that underpins the
public/private partnership that lies at the heart of academic publishing. The
journals worry that their main income stream (allowing them to provide studentships
etc) will be eliminated; while the publishers worry that their privileged
position between subsidised creators of content and subsidised buyers of
content will be squeezed. Both these
anxieties are justified.
But, we need to
ask ourselves whether we really want to use the roundabout and expensive route
of generating income from University Library budgets via the publication of
materials produced by academic staff, to take money from the providers of
education - the Universities - in order to give it to the journals and scholarly
societies, in order to allow them, to in turn purchase education from the
Universities. It is ridiculous.
As for the academic presses, they have spent thirty years
squeezing the 'added value' from their operation. In-house copy-editing and proof reading for
the most part went ages ago. And many
presses now demand what amounts to 'camera ready' copy. If the presses do not want to serve their
traditional role in an ecology of scholarship (sifting and polishing its
products), then it is not clear what their profits are based on. At the moment, the greatest input on the
part of the presses lies in advertising and licensing content, policing its
re-use and in producing hard copy versions of books and articles that are largely
unwanted (ask any librarian). A
thoroughgoing Open Access model eliminates the need for selling, licensing, and
policing, while time will take care of the romantic attachment to wood pulp.
Current debate seems most fully motivated by a reactionary
and defensive fear that a change in the nature of academic publication will
unravel the systems of authority and organisational finance that used to
deliver public debate. But, if we have
faith in the importance of the academy and of scholarship, then we need to
continually re-invent the process. Open
Access provides a perfect opportunity to reconnect with the founding principles
of the academy.
Sunday, 23 October 2011
Academic History Writing and its Disconnects
This is the rough text of a short talk I am scheduled to deliver at a symposium on 'Future Directions in Book History' at Cambrdige on the 24th of November 2011.
I am on the programme as talking briefly about the ‘OldBailey Online and other resources’ (by which I assume is meant London Lives,
Connected Histories, and Locating London’s Past, and the other websites I have
helped to create over the last ten or twelve years). But I am afraid I
have no interest whatsoever in discussing the Old Bailey or the other
websites. The hard intellectual work
that went in to their creation was done between 1999 and 2010, and for the most
part they have found an audience and a user base and will have their own
impact, without me having to discuss them any further. We know how to do this stuff, and anyone can
read the technical literature, and I very much encourage you to do so.
Instead, I want to talk about how the evolution of the forms
of delivery and analysis of text inherent in the creation of the online,
problematizes and historicises the notion of the book as an object, and as a
technology; and in the process problematizes the discipline of history itself as we practise it in the digital present.
The project of putting billions of words of keyword
searchable stuff out there is now nearing completion. We are within sight
of that moment when all printed text produced between 1455 and 1923 (when the Disney Corporation has determined that the needs of modern corporate capitalism trumped the Enlightenment ideal), will be available
online for you to search and read. The
vast majority of that text is currently configured to pretend to be made up of
‘books’ and other print artefacts, But,
of course, it is not. At some level it
is just text – the difference between one book and the next a single line of
metadata. The hard leather covers that
used to divide one group of words from another are gone; and every time you
choose to sit comfortably in your office reading a screen, instead of going to
a library or an archive, while kidding yourself that you are still reading a ‘book’,
you are in fact participating in a charade.
We are swimming in deracinated, Google-ised, Wikipedia-ised text.
In other words, and let’s face it: the book as a technology
for packaging and delivery, storing and finding text is now redundant. The underpinning mechanics that determined
its shape and form are as antiquated as moveable type. And in the process of moving beyond the book,
we have also abandoned the whole post-enlightenment infrastructure of libraries
and card catalogues (or even OPACS), of concordances, and indexes and tables of
contents. They are all built around the
book, and the book is dead.
If this all sounds rather doom laden and apocalyptic – and
no doubt we could argue about the rosy future and romantic appeal of the hard
copy book – it shouldn’t. At least as
far as the ‘history of the book’ is concerned these developments have been
entirely positive
First, it has allowed us to begin to escape the intellectual
shackles that the book as a form of delivery, imposed upon us. If we can escape the self-delusion that we
are reading ‘books’, the development of the infinite archive, and the creation
of a new technology of distribution, actually allows us to move beyond the linear
and episodic structures the book demands, to something different and more
complex. It also allows us to more
effectively view the book as an historical artefact and now redundant form of
controlling technology. The 'book' is newly
available for analysis.
The absence of books makes their study more important, more
innovative, and more interesting. It
also makes their study much more relevant to the present – a present in which
we are confronted by a new, but equally controlling and limiting technology for
transmitting ideas. By mentally escaping
the ‘book’ as a normal form and format, we can see it more clearly for what it
was. And to this extent, the death of
the book is a fantastic and liberating thing – the fascism of the format is beaten.
At the same time, I think we are confronted by a profound
intellectual challenge that addresses the very nature of the historical
discipline. This transition from the
‘book’, to something new, fundamentally undercuts what we do more generally as
‘historians’. When you start to unpick the nature of the historical
discipline, it is tied up with the technologies of the printed page and the
book in ways that are powerful and determining.
Our footnotes, our post-Rankean cross referencing and practises of
textual analysis are embedded within the technology of the book, and its
library.
Equally, our technology of authority – all the visual and
textual clues that separate a CUP monograph from the irresponsible musings of a
know-nothing prose merchant – are slipping away.
While our professional identity – the titles, positions and honorifics – built
again on the supposedly secure foundations of book publishing – is ever less compelling. So the question then becomes, is history – particularly in
its post-Rankean, professional and academic form - dead? Are we losing that beautiful disciplinary
character that allows us to think beyond the surface, and makes possible complex analyses that transcend mere cleverness?
And on the face of it, the answer is yes – the renewed role
of the popular block buster, and an every growing and insecure emphasis on
readership over scholarship, would suggest that it is. In Britain we shy away from the
metrics that would demonstrate ‘impact’ primarily because we fear that we may not have any.
Collectively we have put our heads in the sands, and our
arses in the air, and seemingly invited the world to take a shot. A single and self-evident instance that
evidences a deeper malaise is our current failure to bother citing
what we read. We read
online journal articles, but cite the hard copy edition; we do keywords
searches, while pretending to undertake immersive reading. We search 'Google Books', and pretend we are not.
But even more importantly, we ignore the critical impact of
digitisation on our intellectual praxis.
Only 48% of the significant words in the Burney collection ofeighteenth-century newspapers are correctly transcribed as a result of poor OCR. This makes the other 52% completely
un-findable. And of course, from the
perspective of the relationship between scholarship and sources, it is always
the same 52%. My colleague Bill Turkel,
describes this as the Las Vegas effect – all bright lights, and an invitation
to instant scholarly riches, but with no indication of the odds, and no exit
signs. We use the Burney collection
regardless – not even bothering to apply the kind of critical approach that
historians have built their professional authority upon. This is roulette dressed up as scholarship.
In other words, we have abandoned the rigour of traditional
scholarship. Provenance, edition,
transcription, editorial practise, readership, authorship, reception – the things we query issues in relation to books, are left unexplored in relation to the online text we actually read.
And as importantly, the way we promulgate our ‘history’ has
not kept up either. I want television
programmes with footnotes, and graphs with underlying spreadsheets and
sliders. Yes, I want narrative and
analysis, structure, point and purpose.
I want to continue to be able to engage in the grand conversation that
is history; but it cannot continue to be produced as a ragged and impotent ghost
of a fifteenth century technology; and if we don’t do something about it, we
might as well all go off and figure out how to write titillating tales of
eighteenth-century sex scandals, because at least they sell.
The book had a wonderful 1200 odd year history, which is
certainly worth exploring. Its form self-evidently controlled and informed significant aspects of
cultural and intellectual change in the West (and through the impositions of
Empire, the rest of the world as well); but if, as historians, we are to avoid
going the way of the book, we need to separate out what we think history is
designed to achieve, and to create a scholarly technology that delivers it.
In a rather intemperate attack on the work of Jane Jacobs,
published in 1962, Louis Mumford observed that:
‘… minds unduly fascinated by
computers carefully confine themselves to asking only the kind of question that
computers can answer and are completely negligent of the human contents or the human results.’
LewisMumford, “The Sky Line "Mother Jacobs Home Remedies",” The New Yorker, December 1, 1962, p. 148
I am afraid that in the last couple of decades, historians who are unduly fascinated by books, have restricted themselves to asking only the kind of questions books can answer. Fifty years is a long time in computer science. It is about time we found out if a critical and self-consciously scholarly engagement with computers might not now allow us to more effectively address the ‘human contents’ of the past.
Subscribe to:
Posts (Atom)