I have just spent a couple of days at the annual conference of the NFAIS (National Federation of Advanced Information Services) in Philadelphia. I was giving a brief paper at the request of the good people at the British Library about developments in text mining in the Humanities, and I was happy to be invited and to participate.
But it was only when I had sat through a day or two of presentations that I realised just how out of place I was, and how irrelevant my comments were. It turns out the NFAIS is the trade organisation for all the companies (and some libraries) that have been building a commercial operation for a hundred years by placing themselves between information and those who need it. Thompson-Reuters were heavily represented, as was Cengage/Gale, several medical abstracting services, and a host of companies providing data to particular sectors of the economy such as the building trades and architects. And they were all presenting their well articulated models of data gathering and manipulation designed to deliver a pablum of stuff to the desktops of America's commercial movers, shakers and capitalists. Interestingly, the people who weren't there, were Google, Facebook, Twitter, the Creative Commons or representatives of the Open Access movement. Neither new model capitalists, nor Open Access evangelists were present. And while there was a strand of discussion focussed on research and academic library services, this was a small corner of an essentially old style commercial ecology. Nor were the real innovations in data modelling and analysis coming out of CS on display. A constant sub-theme of the conference seemed to be a tetchy criticism of Google for having done a half-arsed job of inter-mediating between data and users, while the Twitter stream for this event was almost non-existent. It was clear that these data professionals were having their conversation somewhere else - though I never did find out where.
There was a lot of talk about Altmetrics as a way of adding value to the data people already had, prior to selling it to the managers of research and education. And the theme of the event appeared to be a call to extend a hundred year old business model to ensure that these companies were delivering precisely the data that people needed (rather than what they thought they wanted), in a form that allowed them to use that data without thinking. The controlling metaphor was - people buy a drill, because they really want a hole.
I was bemused by this. I don't want a hole. I want a drill, a hammer, a saw and a workshop to make stuff in. And I certainly don't want anyone else to second guess what it is I am making (perhaps wonky, but original).
And then it occurred to me where my disconnect came from. The NFAIS and the companies they represent derive from a long and largely American tradition of late Enlightenment data processing. Their origins lie in Union Catalogues and abstracting services via microfilm; in creating a pre-digested, post-Enlightenment world of understood data, that could be packaged and catalogued and sold as yard after yard of uniform reference volumes. NFAIS used to stand for the National Federation of Abstracting and Information Services. I have always put the American obsession with this kind of thing down to its inability to get over the European Enlightenment long after the rest of us got bored.
My overwhelming impression was that all these companies were anxious to widen the gap between data and its users, to ensure that they continued to have a role and an income - tapping the stream between the two for annual profit. Some of the work was reasonably sophisticated (though most of it felt more 'relational database' than anything more innovative), and it was clear that many saw the way forward as providing faster access to real-time data in a form that would become normalised in a business context (or well funded, close to market STEM).
In retrospect, and after having heard the presentations which came after my own, what I really should have said more forcefully, is get out of the way, this is boring, and it misses the point entirely. We are rapidly approaching the stage when the devil's contract between private companies and the public sector, which has governed data delivery in both the humanities and STEM for the last fifteen years on line (and a hundred years off-line) is going to break down. Open Access, for example, is just a wedge issue for a wider re-thinking of how research and data, and its users will interact. And the fact that both the British and American governments were there first, is an indication that this particular community is not paying sufficient attention.
I am not overly exercised by the profiteering of these companies (if people want to sell their souls for a health plan and a cheap suit, that is OK by me). Nor do I really want to castigate them for their place in an information food chain. Companies like Cengage/Gale have coralled a useful amount of money for data processing. I am just struck by the lack of serious engagement with the real changes in the fundamental relationship between the production and consumption of data that the last fifteen years has wrought. More than anything else, my three days in Phili has brought home to me that I simply don't inhabit the same data universe that I did fifteen years ago.
This blog is a space for me to rant in that most seventeenth-century sense of the word; and to cut and paste the ideas and comments that don't seem to fit in more traditional forms of academic publication.
Showing posts with label librarianship. Show all posts
Showing posts with label librarianship. Show all posts
Wednesday, 27 February 2013
Monday, 29 October 2012
A Five Minute Rant for the Consortium of European Research Libraries
I have been asked to participate in a panel at the annual CERL conference - and to speak for no more than five minutes or so. Initially, I was just going to wing it, but then, in writing up a couple notes, five minutes worth of text found its way on to the screen. In the spirit of never wasting a grammatical sentence, the text is below. I probably wont follow it at the conference, but it reflects what I wanted to say.
CERL - British Library, 31 October 2012:
We all know just how transformative
the digitisation of the inherited print archive has been. Between Google Books, ECCO, EEBO, Project
Gutenberg, the Burney Collection, the British Library's 19th century
newspapers, the Parliamentary Papers, the Old Bailey Online, and on and on,
something new has been created. And it
is a testament to twenty years of seriously hard graft. But, it is one of the great ironies of the
minute that the most revolutionary technical change in the history of the human
ordering of information since writing - the creation of the infinite archive,
with all its disruptive possibilities - has resulted in a markedly conservative
and indeed reactionary model of human culture.
For both technical and legal reasons,
in the rush to the on line, we have given to the oldest of Western canons a new
hyper-availability, and a new authority. With the exception of the genealogical sites,
which themselves reflect the Western bias of their source materials and
audience, the most common sort of historical web resource is dedicated to
posting the musings of some elite, dead, white, western male - some scientist,
or man of letters; or more unusually, some equally elite, dead white woman of
letters. And for legal reasons as much
as anything else, it is now much easier to consult the oldest forms of
humanities scholarship instead of the more recent and fully engaged
varieties. It is easier to access work from
the 1890s, imbued with all the contemporary relevance of the long dead, than it
is to use that of the 1990s.
Without serious intent and political
will - a determination to digitise the more difficult forms of the
non-canonical, the non-Western, the non-elite and the quotidian - the materials
that capture the lives and thoughts of the least powerful in society - we will
have inadvertently turned a major area of scholarship, in to a fossilised
irrelevance.
And this is all the more important
because just at the same moment that we have allowed our cultural inheritance
to be sieved and posted in a narrowly canonical form; the siren voices of the
information scientists: the Googlers, coders and Culturomics wranglers, have
discovered in that body of digitised material, a new object of study. All digital texts is now data, that date is
now available for new forms of analysis, and that data is made up of the stuff
we chose to digitise. All of which embeds a subtle biase towards a
particular subset of the human experience.
Using measures derived from Ngrams, and topic modelling, natural
language processing, and TF-IDF similarity measures; scientists are beginning
to use this text/data as the basis for a new search for mathematically
identifiable patterns. And in the
process, the information scientists are beginning to carve out what is being
presented as 'natural' patterns of change, that turn the products of human
culture into a simple facet of a natural, and scientifically intelligible
world. The only problem with this is
that the analysts undertaking this work are not overly worried by the nature of
the data they are using. For most, the
sheer volume of text makes its selective character irrelevant.
But if we are not careful, we will
see the creation of a new 'naturalisation' of human thought based on the
narrowest sample of the oldest of dead white males. And to this particular audience I just want
to suggest that we need to be much more critical about what it is that we
digitise; what we allow to represent the cultures libraries and collections
stand in for; and that we need to engage more comprehensively and intelligently
with the simple fact that we are in the middle of a selective recreation of
inherited culture.
Subscribe to:
Posts (Atom)