Showing posts with label librarianship. Show all posts
Showing posts with label librarianship. Show all posts

Wednesday, 27 February 2013

Drills and holes

I have just spent a couple of days at the annual conference of the NFAIS (National Federation of Advanced Information Services) in Philadelphia.  I was giving a brief paper at the request of the good people at the British Library about developments in text mining in the Humanities, and I was happy to be invited and to participate.

But it was only when I had sat through a day or two of presentations that I realised just how out of place I was, and how irrelevant my comments were.  It turns out the NFAIS is the trade organisation for all the companies (and some libraries) that have been building a commercial operation for a hundred years by placing themselves between information and those who need it.  Thompson-Reuters were heavily represented, as was Cengage/Gale, several medical abstracting services, and a host of companies providing data to particular sectors of the economy such as the building trades and architects.   And they were all presenting their well articulated models of data gathering and manipulation designed to deliver a pablum of stuff to the desktops of America's commercial movers, shakers and capitalists.  Interestingly, the people who weren't there, were Google, Facebook, Twitter, the Creative Commons or representatives of the Open Access movement.  Neither new model capitalists, nor Open Access evangelists were present.  And while there was a strand of discussion focussed on research and academic library services, this was a small corner of an essentially old style commercial ecology.  Nor were the real innovations in data modelling and analysis coming out of CS on display.  A constant sub-theme of the conference seemed to be a tetchy criticism of Google for having done a half-arsed job of inter-mediating between data and users, while the Twitter stream for this event was almost non-existent.  It was clear that these data professionals were having their conversation somewhere else - though I never did find out where. 

There was a lot of talk about Altmetrics as a way of adding value to the data people already had, prior to selling it to the managers of research and education.  And the theme of the event appeared to be a call to extend a hundred year old business model to ensure that these companies were delivering precisely the data that people needed (rather than what they thought they wanted), in a form that allowed them to use that data without thinking.  The controlling metaphor was - people buy a drill, because they really want a hole.

I was bemused by this.  I don't want a hole.  I want a drill, a hammer, a saw and a workshop to make stuff in.  And I certainly don't want anyone else to second guess what it is I am making (perhaps wonky, but original).

And then it occurred to me where my disconnect came from.  The NFAIS and the companies they represent derive from a long and largely American tradition of late Enlightenment data processing.  Their origins lie in Union Catalogues and abstracting services via microfilm; in creating a pre-digested, post-Enlightenment world of understood data, that could be packaged and catalogued and sold as yard after yard of uniform reference volumes. NFAIS used to stand for the National Federation of Abstracting and Information Services. I have always put the American obsession with this kind of thing down to its inability to get over the European Enlightenment long after the rest of us got bored.

My overwhelming impression was that all these companies were anxious to widen the gap between data and its users, to ensure that they continued to have a role and an income - tapping the stream between the two for annual profit.  Some of the work was reasonably sophisticated (though most of it felt more 'relational database' than anything more innovative), and it was clear that many saw the way forward as providing faster access to real-time data in a form that would become normalised in a business context (or well funded, close to market STEM).  

In retrospect, and after having heard the presentations which came after my own, what I really should have said more forcefully, is get out of the way, this is boring, and it misses the point entirely.  We are rapidly approaching the stage when the devil's contract between private companies and the public sector, which has governed data delivery in both the humanities and STEM for the last fifteen years on line (and a hundred years off-line) is going to break down.  Open Access, for example, is just a wedge issue for a wider re-thinking of how research and data, and its users will interact.  And the fact that both the British and American governments were there first, is an indication that this particular community is not paying sufficient attention.

I am not overly exercised by the profiteering of these companies (if people want to sell their souls for a health plan and a cheap suit, that is OK by me).  Nor do I really want to castigate them for their place in an information food chain.  Companies like Cengage/Gale have coralled a useful amount of money for data processing.  I am just struck by the lack of serious engagement with the real changes in the fundamental relationship between the production and consumption of data that the last fifteen years has wrought.  More than anything else, my three days in Phili has brought home to me that I simply don't inhabit the same data universe that I did fifteen years ago. 






Monday, 29 October 2012

A Five Minute Rant for the Consortium of European Research Libraries



I have been asked to participate in a panel at the annual CERL conference - and to speak for no more than five minutes or so.  Initially, I was just going to wing it, but then, in writing up a couple notes, five minutes worth of text found its way on to the screen.  In the spirit of never wasting a grammatical sentence, the text is below.  I probably wont follow it at the conference, but it reflects what I wanted to say.
CERL - British Library, 31 October 2012:
We all know just how transformative the digitisation of the inherited print archive has been.  Between Google Books, ECCO, EEBO, Project Gutenberg, the Burney Collection, the British Library's 19th century newspapers, the Parliamentary Papers, the Old Bailey Online, and on and on, something new has been created.  And it is a testament to twenty years of seriously hard graft.  But, it is one of the great ironies of the minute that the most revolutionary technical change in the history of the human ordering of information since writing - the creation of the infinite archive, with all its disruptive possibilities - has resulted in a markedly conservative and indeed reactionary model of human culture.

For both technical and legal reasons, in the rush to the on line, we have given to the oldest of Western canons a new hyper-availability, and a new authority.  With the exception of the genealogical sites, which themselves reflect the Western bias of their source materials and audience, the most common sort of historical web resource is dedicated to posting the musings of some elite, dead, white, western male - some scientist, or man of letters; or more unusually, some equally elite, dead white woman of letters.  And for legal reasons as much as anything else, it is now much easier to consult the oldest forms of humanities scholarship instead of the more recent and fully engaged varieties.  It is easier to access work from the 1890s, imbued with all the contemporary relevance of the long dead, than it is to use that of the 1990s.  

Without serious intent and political will - a determination to digitise the more difficult forms of the non-canonical, the non-Western, the non-elite and the quotidian - the materials that capture the lives and thoughts of the least powerful in society - we will have inadvertently turned a major area of scholarship, in to a fossilised irrelevance.

And this is all the more important because just at the same moment that we have allowed our cultural inheritance to be sieved and posted in a narrowly canonical form; the siren voices of the information scientists: the Googlers, coders and Culturomics wranglers, have discovered in that body of digitised material, a new object of study.  All digital texts is now data, that date is now available for new forms of analysis, and that data is made up of the stuff we chose to digitise.   All of which embeds a subtle biase towards a particular subset of the human experience.  Using measures derived from Ngrams, and topic modelling, natural language processing, and TF-IDF similarity measures; scientists are beginning to use this text/data as the basis for a new search for mathematically identifiable patterns.  And in the process, the information scientists are beginning to carve out what is being presented as 'natural' patterns of change, that turn the products of human culture into a simple facet of a natural, and scientifically intelligible world.  The only problem with this is that the analysts undertaking this work are not overly worried by the nature of the data they are using.  For most, the sheer volume of text makes its selective character irrelevant.

But if we are not careful, we will see the creation of a new 'naturalisation' of human thought based on the narrowest sample of the oldest of dead white males.  And to this particular audience I just want to suggest that we need to be much more critical about what it is that we digitise; what we allow to represent the cultures libraries and collections stand in for; and that we need to engage more comprehensively and intelligently with the simple fact that we are in the middle of a selective recreation of inherited culture.