Showing posts with label history. Show all posts
Showing posts with label history. Show all posts

Friday, 29 May 2015

The UK Web Archive, Born Digital Sources and Rethinking the Future of Research

The following post is derived from a short talk I gave at a doctoral training event at the British Library in May 2015, focused on using the UK Web Archive.  It was written with PhD students in mind, but really forms a meditation on the opportunities created when we are working with web sites rather than print.  While lightly edited, the text retains the ticks and repetitions of public presentation.



My office c.1984
I normally work on properly dead people of the sort that do not really appear in the UK Web Archive – most of them eighteenth-century beggars and criminals.  And in many respects the object of study for people like me – interlocutors of the long dead -  has not changed that much in the last twenty years.  For most of us, the ‘object of study’ remains text.  Of course the ‘digital’ and the online has changed the nature of that text.  How we find things – the conundrums of search – shape the questions we ask.  And a series of new conundrums have been added to all the old ones – does, for instance, ‘big data’ and new forms of visualisation, imply a new ‘open eyed’ interrogation of data?  Are we being subtly encouraged to abandon older social science ‘models’, for something new?   And if we are, should these new approaches take the form of ‘scientific’ interrogation, looking for ‘natural’ patterns – following the lead of the Culturomics movement; or perhaps take the form of a re-engagement with the longue durĂ©e– in answer to the pleas of the History Manifesto.   Or perhaps we should be seeking a return to ‘close reading’ combined with a radical contextualisation - looking at the individual word, person, and thing – in its wider context, preserving focus across the spectrum.

And of course, the online and the digital also raises issues about history writing as a genre and form of publication.   Open access, linked data, open data, the 'crisis' of the monograph, and the opportunities of multi-modal forms of publication, all challenge us to think again about the kind of writing we do, as a  literary form.  Why not do your PhD as a graphic novel? Why not insist on publishing the research data with your literary over-lay?  Why not do something different?  Why not self-publish?

These are conundrums all – but conundrums largely of the ‘textual humanities’.  

Ironically, all these conundrums have not had much effect on the academy and the kind of scholarship the academy values.  The world of academic writing is largely, and boringly, the same as it was thirty years ago.  How we do it has changed, but what it looks like feels very familiar.

But the born digital is different.  Arguably, the sorts of things I do, history writing focused on the  properly dead, looks ‘conservative’ because it necessarily engages with the categories of knowing that dominated the nineteenth and twentieth centuries – these were centuries of text, organised into libraries of books, and commentated on by cadres of increasingly professional historians.  The born digital – and most importantly the UK web archive – is just different.  It sings to a different tune, and demands different questions – and if anywhere is going to change practise, it should be here. 

Somewhat to my frustration, I don’t work on the web as an ‘object of study’ –  and therefore feel uncertain about what it can answer and how its form is shaping the conversation; but I did want to suggest that the web itself and more particularly the UK Web Archive provides an opportunity to re-think what is possible, and to rethink what it is we are asking; how we might ask it, and to what purpose.

And I suppose the way I want to frame this is to suggest that the web itself brings on to a single screen, a series of forms of data that can be subject to lots of different forms of analysis.  A few years ago, when APIs were first being advocated as a component of web design, the comment that really struck me, was that the web itself is a form of API, and that by extension the Web Archive is subject to the same kind of ‘re-imagination’ and re-purposing that an API allows for a single site or source.  

As a result, you can – if you want – treat a web page as simple text – and apply all the tools of distant reading of text - that wonderful sense that millions of words can be consumed in a single gulp.   You can apply ‘topic modelling’, and Latent Semantic Analysis; or Word Frequency/Inverse Document Frequency measures.  Or, even more simply; you can count words, and look for outliers – stare hard at the word on the web!

But you can also go well beyond this.  In performance art, in geography and archaeology, in music and linguistics, new forms of reading are emerging with each passing year that seem to me to significantly challenge our sense of the ‘object of study’ – both traditional text and web page.  In part, this is simply a reflection of the fact that all our senses and measures are suddenly open to new forms of analysis and representation.  When everything is digital – when all forms of stuff come to us down a single pipeline -  everything can be read in a new way.  

 Consider for a moment the ‘LIVE’ project from the Royal Veterinary College in London, and their ‘haptic simulator’.  In this instance they have developed a full scale ‘haptic’ representation of a cow in labour, facing a difficult birth, which allows students to physically engage and experience the process of manipulating a calf in situ.  I haven’t had a chance to try this, but I am told that it is a mind altering experience.  It suggests that reading can be different; and should include the haptic - the feel and heft of a thing in your hand.  This is being coded for millions of objects through 3d scanning; but we do not yet have an effective way of incorporating that 3d text into how we read the past. 

 The same could be said of the aural - that weird world of sound on which we continually impose the order of language, music and meaning; but which is in fact a stream of sensations filtered through place and culture.  


Projects like the Virtual St Paul's Cross, which allows you to ‘hear’ John Donne’s sermons from the 1620s, from different vantage points around the yard, changes how we imagine them, and moves from ‘text’ to something much more complex and powerful.  And begins to navigate that normally unbridgeable space between text and the material world.  And if you think about this in relation to music and speech online – you end up with something different on a massive scale.

One of my current projects is to create a sound scape of the courtroom at the Old Bailey - to re-create the aural experience of the defendant - what it felt like to speak to power, and what it felt like to have power spoken at you from the bench. And in turn, to use that knowledge to assess who was more effective in their dealings with the court, and whether, having a bit of shirt to you, for instance, effected your experience of transportation or imprisonment.  And the point of the project is to simply add a few more variables to the ones we can securely derive from text.

It is an attempt to add just a couple of more columns to a spreadsheet of almost infinite categories of knowing.  And you could keep going – weather, sunlight, temperature, the presence of the smells and reeks of other bodies.  Ever more layers to the sense of place.  In part, this is what the gaming industries have been doing from the beginning, but it also becomes possible to turn that creativity on its head, and make it serve a different purpose.

In the work of people such as Ian Gregory, we can see the beginnings of new ways of reading both the landscape, and the textual leavings of dead.  Bob Shoemaker, Matthew Davies and I (with a lot of other people) tried to do something similar with Old Bailey material, and the geography of London in the Locating London’s Past project.

This map is simply colours blue, red and yellow mapped against brown and green.  I have absolutely no idea what this mapping actually means, but it did force me to think differently about the feel and experience of the city.  And I want to be able to do the same for all the text captured in the UK domain name. 

All of which is to state the obvious.  There are lots of new readings that change how we connect with historical evidence – whether that is text, or something more interesting.    In creating new digital forms of inherited culture - the stuff of the dead - we naturally innovate, and naturally enough, discover ever changing readings.  But the Web Archive, challenges us to do a lot more; and to begin to unpick what you might start pulling together from this near infinite archive. 

In other words, the tools of text are there, and arguably moving in the right direction, but there are several more dimensions we can exploit when the object of study is itself an encoding.

Each web page, for instance, embodies a dozen different forms.  Text is obvious, but it is important to remember that each component of the text – each word and letter, on a web page - is itself a complex composite.  What happens when you divide text by font or font size; weight, colour, kerning, formatting etc.  By location - in the header, or the body, or wherever the CSS sends it; or more subtly by where it appears to a users’ eye - in the middle of a line – or at the end.

Suddenly, to all the forms of analysis we have associated with ‘distant reading’ there are five or six further columns in the spread sheet – five or six new variables to investigate in that ‘big data’ eye-opened sort of way.

And that is just the text.  The page itself is both a single image, and a collection of them – each with their own properties.  And one of the great things that is coming out of image research is that we can begin to automate the process of analysing those screens as ‘images’.  Colour, layout, face recognition etc.  Each page, is suddenly ten images in one – all available as a new variable; a new column in the spreadsheet of analysis.  And, of course, the same could be said of embedded audio and video.

And all of that is before we even look under the bonnet.  The code, the links, the meta data for each page – in part we can think of these as just another iteration of the text; but more imaginatively, we can think about it as more variables in the mix.

But, of course, that in itself miss-understands the web and the Web Archive.  The commonplace metaphor I have been using up till now is of a ‘page’ – and is the intellectual equivalent of skeumorphism - relying on material world metaphors to understand the online.

But these aren’t pages at all, they are collections of code and data that generate in to an experience in real time.  They do not exist until they are used - if a website in the forest is never accessed, it does not exists.  The web archive therefore is not an archive of ‘objects’ in the traditional sense, but a snapshot from a moving film of possibilities.  At its most abstract, what the UK Web Archive has done, is spirit in to being the very object it seeks to capture – and of course, we all know that in doing so, the capturing itself changes the object.  Schrödinger's cat may be alive or dead, but its box is definitely open, and we have visited our observations upon its content.

So to add to all the layers of stuff that can fill your spreadsheet, there also needs to be columns for time and use; re-use and republication.  And all this is before we seek to change the metaphor and talk about networks of connections, instead of pages on a website.

Where I end up is seriously jealous of the possibilities; and seriously wondering what the ‘object of study’ might be.  In the nature of an archives, the UK Web Archive imagines itself as an ‘object of study’; created in the service of an imaginary scholar.  The question it raises is how do we turn something we really can’t understand, cannot really capture as an object of study, to serious purpose?  How do we think at one and the same time of the web as alive and dead, as code, text, and image – all in dynamic conversation one with the other.  And even if we can hold all that at once, what is it are we asking?

Friday, 3 January 2014

Judging a book by its URLs

It will sound odd, but I have recently had a great time editing URLs.  Robert Shoemaker and I have have just finished a book for CUP, derived from the London Lives project, and called - London Lives: Poverty, Crime and the Making of a Modern City, 1690-1800. It is a long book (170,000 words) and each quote and reference in it is linked via a URL to the original document or article, book or web-resource used as evidence or to contextualize the argument.  It will be published as both an ebook and in hard copy, and the links need to be robust, and secure.  My estimate is that there are in the region of 4,000 URLs included in the manuscript (which was written collaboratively in PMWiki).  In the end, I found that I could identify an appropriate link for 98% of all footnote references, but then had to eliminate around 10% of these, as the relevant URL was just not useable.  The book took some nine years, and I am glad it is finished.

One of my final jobs was editing those 4000 URLs.   It took about three months work, spread over the last year, and I have just finished spending a week or so confirming what I hope will be their final form.  When I have told people about this work many have looked incredulous and suggested that this is the sort of technical implementation process that should be left to others.  A couple of otherwise nice people have suggested I dump this job on the shoulders of the nearest PhD student.  But for myself, it is precisely the kind of thing that an author should do for themselves.  And in doing it, two things kept coming to mind.  First was how the role of the scholar in creating a rigorous academic apparatus is a central part of the intellectual journey that academic writing involves - and that we should see the implementation of the online version of this in the light of the precise writing of footnotes and references that mark out good scholarship.  And second, that URLs encode a system of design and intent, online architecture and system of access, that signal the quality and permanence (the academic credibility and perceived audience) of historical materials online.  And that just as we have always sorted and judged scholarship by its form, we should think a bit harder about how the form of a URL can let us interrogate online materials.

On the first point, I do not know of much discussion of the joys of this kind of academic slog.  There is a lot of good writing on research and archives (by Carolyn Steedman and Arlette Farge among many others), on writing and thinking, but no-one talks much about the painstaking labour that goes in to turning a rough draft in to a final finished piece of scholarship.  And here I am really talking about generating accurate and fully comprehensive footnotes that reflect both the material cited, and the research journey that resulted in the main text.   This has become much easier with online catalogs and citation management packages, but nevertheless remains laborious and a reflection of our collective and individual commitment to a particular kind of evidenced discussion.  But for me it also represents my favourite compromise.  The writing of history is a wonderfully imaginative and creative process.  And in some respects we wish to judge the product of history writing as art.  Is it enjoyable to read? Is it convincing?  Does it do the job of good writing in liberating the readers' imagination?  In making these judgements we tend to appeal to a notion of 'value' that is cultural and that privileges dominant forms of authority.  This aspect of judgement is essentially romantic; with all the implications for western and elite hegemony embedded in that idea.  At the same time history writing is the result of simple hard work of a more technical kind - in the archives, in collating and collecting, re-ordering and interrogating data.  And it is valuable because it encompasses that hard work.  The beauty of the academic apparatus is that it evidences this and in the process generates a different measure of value.  In other words it is where quality is tied to a 'labour theory of value'.  I love the academic slog because it is where un-moored judgement is tied down to hard labour; and where value can be universalized in a common human experience (work).  In other words I really enjoyed editing 4000 URLs precisely because in them and their associated footnotes lies a claim to and evidence of the hard labour that underpins the book itself.

 At the same time, the process also taught me to read URLs differently.  Clearly coders and web designers do this as a matter of course.  But I am a historian and want to read URLs as a scholar, rather than as a programmer or designer.  And for me, the important thing is that URLs embed the structure of a site, making it plain to see for anyone willing to look hard; and that they are made up of both the character of a library reference, and a command directed at the new technology of discovery - the Internet .  There are just lots of different types of URL.

There are 'Search URLs' that include all the elements that  take the user past a collection to a specific object, but don't let you go directly there without the query.  And there are URLs that encode a cataloging hierarchy.  There are URLs that sift data, or work in your browser to change the data delivered, highlighting phrases or sifting material.  And there are URLs that encode licensing, passwords, and access information.  It is easy enough to find that the whole search journey that took you from a library catalog to an individual item is encoded directly in the URL, and even personalized to you, the machine you are using, or the forms of access you can deploy.  It is easy to find URLs that run on for hundreds of characters, each element divided by a '&' or a '%', or such.

But in creating robust reproducible links to credible historical materials most of these URLs are at least problematic if not useless.  If they include details for institutional access, or session information, they cannot be re-used by someone else.  These URLs are friable and fragile things and not fit for scholarly purposes.  And as a result, for the London Lives book we have been forced to eliminate all the links we originally hoped to include to forty or fifty different sites.  To take a single example, most archives structure their online collections with search in mind, making it difficult to link to a single item.  I spent a lot of time finding the catalog entry for every manuscript we cited in the London Metropolitan Archives, and Westminster Archives Centre, only to regretfully strip out the links when confronted by a complex URL that just did not look credible as a long term citation of the item itself.

Even in its simplest, and in the form recommended by the site for sharing a link, a London Metropolitan Archives URL looks like this:

http://search.lma.gov.uk/scripts/mwimain.dll/144/LMA_OPAC/web_detail/REFD+P69~2FBRI~2FB~2F001~2FMS06554~2F004?SESSIONSEARCH


Since we had consulted these items in their physical form in any case, it did not seem too problematic to leave out these links, but a shame nevertheless.  And likewise, with paywall material there seemed little point in dangling real access, and the promise of credible evidence, before the eyes of readers who would not be able to go beyond the login screen.  It seemed better to cite a specific item in combination with a general (unlinked) URL and date of consultation as reflecting our own research journey, rather than to promise access when we could not deliver it.

With few exceptions the URLs that have been retained (and there are still 4000 of them) address specific items with a specific ID, and usually run to 20 to 40 characters.  DOIs are not bad once you figure out their structure and reformulate them as they should be, rather than the way they are normally cited on journal web pages.

dx.doi.org/10.1353/sec.2010.0268

And Google Books creates a very nice URL once you strip out all the complex formatting instructions that are normally generated as part of a search and inserted after the main ID.  This is what a Google Books' URL looks like if you were to use the 'search' version:

 http://books.google.co.uk/books?id=1sMJGt7_rTAC&printsec=frontcover&dq=%22Prosecution+and+Punishment:+Petty+Crime+and+the+Law%22&hl=en&sa=X&ei=rrzGUq_aDsSy7Aa_9YGQCg&redir_esc=y#v=onepage&q=%22Prosecution%20and%20Punishment%3A%20Petty%20Crime%20and%20the%20Law%22&f=false

And this URL will take to the same book:

 books.google.co.uk/books?id=1sMJGt7_rTAC

 And the Eighteenth-century Short Title Catalog generates some of the most elegant URLs I have found:

estc.bl.uk/T174945

And to a lesser extent, so does the Ethos collection of doctoral theses at the British Library.

ethos.bl.uk/OrderDetails.do?uin=uk.bl.ethos.354762

And London Lives and the Old Bailey Online do pretty well on this score:

www.londonlives.org/browse.jsp?div=LMSMPS501980014
http://www.oldbaileyonline.org/browse.jsp?ref=t17910413-19


In part, I suspect that these issues would all disappear if I had a better sense of the layer of structure that lies beneath the WWW.  But for the moment I am keen to have a short, human-readable URL that looks like it will last longer than the session I am currently logged on for.   All of which simply takes me back to the joy of academic slogging and the importance of the academic apparatus as something that evidences hard work and opens up scholarship to credible criticism that goes beyond simple romantic appreciation and prejudice.

I know all too well that one of the skills of an academic is the ability to judge a book by its cover and the form of the text it contains.   For the online we need to embed URLs into precisely this process - and the joy of all that editing was that at the end of it, I feel I have learned to do just that.



Wednesday, 22 May 2013

Stuff and Dead People

In recent years I find myself using the terms Stuff and Dead People in talks and titles more and more.  And as a historian I find myself conceptualising my work as being about Stuff inherited from Dead People.  Both expressions just sound right.  But it occurs to me that while I have a relatively clear sense of what I am intending to convey when I use these terms, their meanings might not be entirely apparent to others.  For this reason I thought I would have a stab at providing a couple of definitions, and a brief explanation of why I find these terms so useful

In my usage Stuff encompasses all the different varieties of artefact that can be used in practising history.  The term is in some respects an attempt avoid saying that our object of study is text or image, the manmade landscape or a piece of furniture, or indeed even data in its broadest form.  Instead, the use of Stuff is intended to signify that my practise as a historian actively seeks to make use of all of these things.  In terms of an epistemology, it is an attempt to distance myself from the categories of knowing that I (we) have inherited.  Stuff denies the taxonomies of knowing that define a museum object as being different to a pamphlet; a hedgerow different to a  teapot.  In part this usage reflects a profound disillusion with the narrow practise of textual comparison that lies at the heart of the Rankean tradition of historical analysis; but it is also a recognition that new technologies allow us to encompass new types of evidence in new ways.  When all Stuff is data it can be interrogated across boundaries that seemed natural and unbreachable just a few decades ago (between a hedgerow and a teapot). And while data itself is also a form of Stuff, and the transition from varieties of stuff to data is itself a process of creating a new taxonomy, there remains a rather wonderful transition involved.  There is an opportunity to rethink the meanings of Stuff, and without a new vocabulary it is all that much more difficult to do so.

In other words, Stuff is a simple rejection of post-enlightenment categorisation.

In some respects Dead People serves a similar function.  The use of Dead People avoids the traps of both identity and social modelling; while at the same time giving some shape to the object of historical study (human culture in the past).  Ironically some Dead People are still alive.  Henry Kissinger is apparently still breathing, but is nevertheless a figure of substantial historical analysis.  In my view he is undeniably Dead People.  At the same time, because cultural history seems to take longer to turn journalism in to books, Amy Winehouse and Michael Jackson may be dead, but they are not yet Dead People.

The term Dead People implies a refusal to describe the people of past as men or women, workers or citizens, artists or authors.  And in doing so, like Stuff, is used to signal that I do not find the traditional categories and boundaries that comprise social science very helpful.

Stuff we inherit from Dead People is my object of study as a historian. 

One could convey these ideas using other words.  The results might be a bit long winded, but could certainly point up my intention.  At the same time, the use of these terms serve a slightly wider function.  They form an attempt to de-centre the language of historical and social science authority that underpins the professional claims of academic historians as a whole.  By refusing to use the categories and languages of authority we inherited, I am self-consciously rejecting the systems that underpin the professional academic practise of history. 

It is perhaps a ridiculous comparison, but I like to think of the use of these terms as akin to the transition in thinking brought about by the evolution of labelling in quantum theory between the proposal of the eight-fold-way in the 1950s and the November Revolution of 1974.  Like most people of my generation and education, I was raised in an Einsteinian universe in which unusual phenomenon were described in the most secure of scientific jargon - we believed in the physics because it was expressed in the language of authority.  But in the 1970s, in particular, a whole new language of strangeness and charm was broadcast to a popular audience.  As a teenager schooled in an older tradition, this challenged me to rethink.  By using everyday words to describe complex phenomena I was forced to interrogate what I believed more closely than I would otherwise have done.  I don't understand quantum, but suspect I understand Einsteinian physics better as a result!   I use the terms Stuff and Dead People in the hope that their use will challenge listeners to question the labels and phenomena they think I am talking about.
    
    




 

Sunday, 23 October 2011

Academic History Writing and its Disconnects


This is the rough text of a short talk I am scheduled to deliver at a symposium on 'Future Directions in Book History'  at Cambrdige on the 24th of November 2011.


I am on the programme as talking briefly about the ‘OldBailey Online and other resources’ (by which I assume is meant London Lives, Connected Histories, and Locating London’s Past, and the other websites I have helped to create over the last ten or twelve years).  But I am afraid I have no interest whatsoever in discussing the Old Bailey or the other websites.  The hard intellectual work that went in to their creation was done between 1999 and 2010, and for the most part they have found an audience and a user base and will have their own impact, without me having to discuss them any further.  We know how to do this stuff, and anyone can read the technical literature, and I very much encourage you to do so.


Instead, I want to talk about how the evolution of the forms of delivery and analysis of text inherent in the creation of the online, problematizes and historicises the notion of the book as an object, and as a technology; and in the process problematizes the discipline of history itself as we practise it in the digital present. 


The project of putting billions of words of keyword searchable stuff out there is now nearing completion.  We are within sight of that moment when all printed text produced between 1455 and 1923 (when the Disney Corporation has determined that the needs of modern corporate capitalism trumped the Enlightenment ideal), will be available online for you to search and read.  The vast majority of that text is currently configured to pretend to be made up of ‘books’ and other print artefacts,   But, of course, it is not.  At some level it is just text – the difference between one book and the next a single line of metadata.  The hard leather covers that used to divide one group of words from another are gone; and every time you choose to sit comfortably in your office reading a screen, instead of going to a library or an archive, while kidding yourself that you are still reading a ‘book’, you are in fact participating in a charade.  We are swimming in deracinated, Google-ised, Wikipedia-ised text.


In other words, and let’s face it: the book as a technology for packaging and delivery, storing and finding text is now redundant.  The underpinning mechanics that determined its shape and form are as antiquated as moveable type.  And in the process of moving beyond the book, we have also abandoned the whole post-enlightenment infrastructure of libraries and card catalogues (or even OPACS), of concordances, and indexes and tables of contents.  They are all built around the book, and the book is dead. 


If this all sounds rather doom laden and apocalyptic – and no doubt we could argue about the rosy future and romantic appeal of the hard copy book – it shouldn’t.  At least as far as the ‘history of the book’ is concerned these developments have been entirely positive

First, it has allowed us to begin to escape the intellectual shackles that the book as a form of delivery, imposed upon us.  If we can escape the self-delusion that we are reading ‘books’, the development of the infinite archive, and the creation of a new technology of distribution,  actually allows us to move beyond the linear and episodic structures the book demands, to something different and more complex.  It also allows us to more effectively view the book as an historical artefact and now redundant form of controlling technology.  The 'book' is newly available for analysis.


The absence of books makes their study more important, more innovative, and more interesting.  It also makes their study much more relevant to the present – a present in which we are confronted by a new, but equally controlling and limiting technology for transmitting ideas.  By mentally escaping the ‘book’ as a normal form and format, we can see it more clearly for what it was.  And to this extent, the death of the book is a fantastic and liberating thing – the fascism of the format is beaten.


At the same time, I think we are confronted by a profound intellectual challenge that addresses the very nature of the historical discipline.  This transition from the ‘book’, to something new, fundamentally undercuts what we do more generally as ‘historians’.  When you start to unpick the nature of the historical discipline, it is tied up with the technologies of the printed page and the book in ways that are powerful and determining.  Our footnotes, our post-Rankean cross referencing and practises of textual analysis are embedded within the technology of the book, and its library.


Equally, our technology of authority – all the visual and textual clues that separate a CUP monograph from the irresponsible musings of a know-nothing prose merchant – are slipping away.  While our professional identity – the titles, positions and honorifics – built again on the supposedly secure foundations of book publishing – is ever less compelling. So the question then becomes, is history – particularly in its post-Rankean, professional and academic form - dead?  Are we losing that beautiful disciplinary character that allows us to think beyond the surface, and makes possible complex analyses that transcend mere cleverness?
 

And on the face of it, the answer is yes – the renewed role of the popular block buster, and an every growing and insecure emphasis on readership over scholarship, would suggest that it is. In Britain we shy away from the metrics that would demonstrate ‘impact’ primarily because we  fear that we may not have any.


Collectively we have put our heads in the sands, and our arses in the air, and seemingly invited the world to take a shot.  A single and self-evident instance that evidences a deeper malaise is our current failure to bother citing what we read.  We read online journal articles, but cite the hard copy edition; we do keywords searches, while pretending to undertake immersive reading. We search 'Google Books', and pretend we are not.


But even more importantly, we ignore the critical impact of digitisation on our intellectual praxis.  Only 48% of the significant words in the Burney collection ofeighteenth-century newspapers are correctly transcribed as a result of poor OCR.  This makes the other 52% completely un-findable.  And of course, from the perspective of the relationship between scholarship and sources, it is always the same 52%.  My colleague Bill Turkel, describes this as the Las Vegas effect – all bright lights, and an invitation to instant scholarly riches, but with no indication of the odds, and no exit signs.  We use the Burney collection regardless – not even bothering to apply the kind of critical approach that historians have built their professional authority upon.  This is roulette dressed up as scholarship.

In other words, we have abandoned the rigour of traditional scholarship.  Provenance, edition, transcription, editorial practise, readership, authorship, reception – the things we query issues in relation to  books, are left unexplored in relation to the online text we actually read.

And as importantly, the way we promulgate our ‘history’ has not kept up either.  I want television programmes with footnotes, and graphs with underlying spreadsheets and sliders.  Yes, I want narrative and analysis, structure, point and purpose.  I want to continue to be able to engage in the grand conversation that is history; but it cannot continue to be produced as a ragged and impotent ghost of a fifteenth century technology; and if we don’t do something about it, we might as well all go off and figure out how to write titillating tales of eighteenth-century sex scandals, because at least they sell.

The book had a wonderful 1200 odd year history, which is certainly worth exploring.  Its form self-evidently controlled and informed significant aspects of cultural and intellectual change in the West (and through the impositions of Empire, the rest of the world as well); but if, as historians, we are to avoid going the way of the book, we need to separate out what we think history is designed to achieve, and to create a scholarly technology that delivers it.


In a rather intemperate attack on the work of Jane Jacobs, published in 1962, Louis Mumford observed that:


‘… minds unduly fascinated by computers carefully confine themselves to asking only the kind of question that computers can answer and are completely negligent of the human contents  or the human results.’

I am afraid that in the last couple of decades, historians who are unduly fascinated by books, have restricted themselves to asking only the kind of questions books can answer.  Fifty years is a long time in computer science.  It is about time we found out if a critical and self-consciously scholarly engagement with computers might not now allow us to more effectively address the ‘human contents’ of the past.