Blog

  • Back to the whiteboard for inspiration

    Paul Miller (of Talis) was kind enough point out that Talis already has a product named Alto.

    In fact, it’s their ILS.

    D’oh.

    Given that Talis’ corporate name is “composer-based” (and, I don’t know, that they came up with the name waaaaaaaaaaaaaaaaaaaay before I did) I suppose I can relinquish the name “Alto” 🙂

    So, uh, any suggestions for an appropriate name… send ’em to me!

  • Envisioning Alto

    I have mentioned here several times the “alternative to the catalog” project I am trying to implement at Tech. One of the problems that I’ve had is naming the project something that lets people realize what I’m talking about, without the political hairiness of saying “catalog replacement” (since that’s technically not true, anyway).

    In a meeting two weeks ago (about subject guides), I was drawing the concept of this project on the whiteboard of our conference room. It’s been up ever since and in the middle, I had written “ALTOpac” because that was an easy way to loosely describe it in a way that the uninitiated in the room could envision where I was starting from. Sitting in another meeting today, the capitalized letters jumped out at me: ALTO. It means nothing.

    And I like that. Of course it still doesn’t explain what it’s about. That’s what subtitles are for.

    Now, let me explain what the hell Alto is and what it is supposed to do.

    Alto is a “community-based collection builder and search engine”.

    Come to think of it, that might not actually clear anything up.

    Let’s back up a bit, shall we?

    To say searching the catalog is “searching our collection” is quite arbitrary and false. Metasearch doesn’t really solve this problem, since you’d still only point the metasearch engine at certain assets and it’s non-trivial to make relationships between assets. Metasearch is part of the solution, but hardly the panacea.

    Again, our “collection” is an ambiguous term and shouldn’t be solely determined by our collection development policies/budget. It is our opinion that if something is important enough to be added to a reserves list (even a web page), it should technically be part of our collection. I would not, however, say it should be cataloged (and that’s why this isn’t a catalog replacement project, see?). If an item is even bookmarked (via a local social bookmarking service, such as unalog or connotea) it should then become part of our collection. A 1927 engineering textbook from Purdue’s catalog? Index it! If a member of our community finds it important enough to want to come back to and share with a group, it’s important enough for us to aggregate into our “collection”. Relevance comes later (keep reading, if you’re interested).

    There are also relationships that our community (for the sake of argument, let’s start with “Georgia Tech”) builds that are highly relevant for finding connections between disparate “things”. So, the items put on reserve for a particular course have an umbrella of commonality between them that should be utilized for anyone that runs across any of these items. The relevance ranking should be even greater for a user that happens to be a member of the group in question (for instance, is enrolled in the class).

    If Alto has a citation management-esque feature in it, users can very specifically group relevant resources together based on a project. Resources can be anything: books, websites, articles, searches, chat transcripts, trails, you name it.

    And all of this should feed the “relevance beast”, as it were.

    So that’s some background. Given that we’ll have some formal subject classifications for these objects (from the OPAC or from metasearch or whatever), we should be able to bridge the formal to folksonomy to make sense of how people have classified their saved things.

    We can then begin to cluster search results. Format, subject, concept, group, policies… All of these can be browsed after the search begins. The search results will be a combination of metadata objects and library content. If some of the results appear in a given “subject guide”, the guide will a suggested resource (and will, in turn, push some resources into the result set).

    The goal is to open the silos we have created around our resources/services. It would break down the ambiguity between “collections”, “services” and “policies” since they’re all interrelated.

    How do we plan to do this? Glad you asked (you’re still reading, right?)!

    We’ve exported all of the bib records from our catalog. The plan is to use METS as our wrapper around MODS. We’ll then harvest our institutional repository and index our website. That’s a pretty good base to start with. All of this is stored in a dbXML database and indexed with Lucene.

    If users want to harvest a collection from citeseer or OAIster, that will be available and will become part of our collection. Annotations, links to reviews, links to content to index will all be made available.

    I’m leaving a lot out and glossing some of this over… but it starts to put the idea on “paper” for me to come back later.

  • I sound my barbaric YAWP over the walls of my cubicle.

    I woke up at 4:30 this morning.

    One could easily write this off to a variety of stresses: an article I have no business writing; a conference I have no business helping organize; a huge project that I am having problems getting started on; a house that I apparently haven’t sunk enough money in to move into yet; a house that I can’t drag far enough away from the railroad tracks to sell; the usual burden that is “the holidays”… sure one could try to pin it on any of those.

    But I woke up thinking about (meaning that I was dreaming about) something I read recently from Richard Wallis on Panlibus, Talis’ ‘blog:

    Well yes, the current generation of ILS systems were not built with Web Services everywhere. To put it bluntly, who will pay the salaries of the developers who are going to develop these services for you to consume?

    Strange thing to dream about, I know. However, when I think about this one quote, it pisses me off to no end. The University System of Georgia pays Endeavor over $500,000 a year for the privilege of running an ILS that they haven’t invested any innovation in years. Granted, we are 35 libraries, so it’s not like we’re all paying that ransom, but, on the flip side, we’re probably also getting a discount for the very fact that we are so large.

    Then, to think we are but a percentage of Endeavor’s total customer base…

    WHERE IS THAT MONEY GOING, RICHARD?

    Of course, I realize that Talis is in no way related to Endeavor, but I cannot imagine their pricing is so radically different that their coffers have no shillings to pay for developers.

    Besides, they must already have developers, right? Maybe you need hire developers with vision.

    So, to this argument, I call bullshit.

    The other thing that struck me (again, apparently in my dream) is the apologetic tone I see quite frequently (recent example here, lots of others floating about) that shifts the blame of our stagnant, crappy Integrated Library Systems to us, the customers instead to our vendors. The argument goes that we, the libraries, have asked for the wrong things for the ILS and the poor vendors (poor, poor vendors) had their hands tied, literally tied, trying to keep up with our demands to be able to incorporate any sort of innovation in the last 15-20 years. Besides, they’d say, if they came up with something different, libraries might not want the change.

    What (successful) technology company has ever relied on RFPs for their innovation? Are Google’s hands tied until some customer says, “Hey, can you make a web based ‘maps’ site? You know what we need? A new way to do threaded email.”? How about Intel? Microsoft?

    No. These companies realize that they need to innovate to survive. To stagnate or half-ass is the kiss of death. See Novell. For a more dramatic example, see Apple.

    No, it’s time we stop taking it like abused spouses from our vendors. You know, maybe we did overcook the porkchop and maybe we do open our mouths too much, but that’s no reason to have a black eye. If a handful of the better funded libraries were to help found something like the Apache Foundation for library software, our abusive husbands might find treat as partners rather than punching bags. I think I might know a good place to look for talent.

    (In truth, our rottweiler woke me up, but the dream still stands).

  • Library 1.7.02-4 pre 6

    I really, really hate this Library 2.0 meme for a couple of reasons.

    1) All of our problems will not, in fact, be solved with AJAX and web interfaces

    2) In fact many of our problems cannot be solved by technology at all (try doing interesting and meaningful and different work with the current body of MARC records out there and see what I mean)

    3) This quest for 2.0 would be better served if “2.0” was a milestone on the journey to “Library 4.5” — I mean, come on folks, let’s get back into innovating.

    4) I think it trivializes some actually exciting and useful work that I fear will continue to fly under the radar because it’s not “Web 2.0” enough.

    Maybe hype is necessary to rally the troops, but I really wish vision would get more attention.

  • echo(strtolower(COinS));

    Er, anybody object to making COinS all lowercase?

    Because I’m starting to think its case structure is awkward and silly.

  • As much as I liked gravy on my french fries…

    Current weather conditions: 76 degrees and sunny.

  • Reactions from Access 2005

    I’ll have to keep this rather short, since the hotel wireless network is being flaky (as usual) (in fact I am having to write this while standing in the bathroom — oddly the best wireless reception in my room).

    Again, the conference organizers have proven why this is the only professional event I schedule in my year. I’ll comment on the earlier days a little later, but while day 3 is still rather fresh in my mind (as fresh as my poor, oversaturated mind can be), I’d like to touch on a few things.

    1) Listen to every word Lorcan says, always. It blows my mind what an amazing asset he is to our community and how there are people (in my library) who have no idea who he is. His presentation this morning has completely energized me to kick it up a notch in getting our library more into our (and other) user’s “LifeFlow”.

    2) Art and Peter proved that WAG needs no “G”. I’ve been working with Art on this WebDav/OPAC project for months (thanks to SUDOC, as he pointed out), and I never, ever would have dreamed of the things these two are coming up with. Cocoon is, evidentally, a very magical beast and the potential of storing these “trails” could have huge implications on the collaborative research environment that we are trying to create at Georgia Tech. Being able to chart the path of scholarship would make it easier to get to the giant to stand upon his shoulders.

    3) Internet communication is lousy when trying to develop a new spec. Despite being there at the beginning (and being a very loud proponent of COinS), I could not wrap my head around the use cases for COinS-PMH. Oh, Dan tried to “learn me”, but it really took his presentation today to “get it”. I definitely “get it” now, and expect to see COinS-PMH all over Tech.

    4) Hackfest is the greatest invention ever. And I honestly couldn’t imagine it working properly at any other conference (sorry, LITA).

    5) This is why I’m applying to enroll in the Master’s program for Human-Computer Interaction at Tech. This, coupled with the previous two presentations (Art/Peter’s, Dan’s)… Holy crap. The world would be so different.

    6)
    The U.S. is screwed. We have sold our souls, culture and future to corporate interests and I’m not sure how we can fix it. As Peter remarked to me, hopefully Cliff Lynch’s vision of a world where everything is digitized except the intellectual output after 1920 will light a fire under us. I fear at that point it may be too late, however. It looks like Canada’s future might be a bit brighter. Even if it isn’t, I’ll get fired up by the revolutionary rhetoric, any day.

    Wow, I love this place.

  • But, you see, my wheel is nothing like all of those other seemingly identical wheels

    I am still feeling my way around Python. I have yet to grasp the zen of being Pythonic, but I am at least coming to grips with real object orientation (as opposed to the named hashes of PHP) and am actually taking the leap into error handling, which, if you have dealt with any of the myriad bugs in any of my other projects, you’d know has been a bit of a foreign concept to me.

    Python project #2 is known as RepoMan (thanks to Ed Summers for the name). It attempts to solve a problem that not one but two other opensource projects already have solved admirably (I’ll go into more about this in a bit). RepoMan is an OAI Repository indexer that makes said repository available via SRU. I created it in an attempt to make our DSpace implementation searchable from remote applications (namely, the site search and the upcoming alternative opac). It’s an extremely simple two script project that has only taken a week to get running largely due to the existence of two similar and available python scripts that I could modify for my own use. It’s also due to the help of Ed Summers and Aaron Lav.

    The harvester is, basically, Thom Hickey’s one page OAI harvester with some minor modification. I have added error handling (the two lines I added to compensate for malformed xml must have been over the “one page limit”) and instead of outputting to a text file, it shoves the records in a Lucene index (thanks to PyLucene). This part still needs some work (I’m not sure what it would do with an “updated” record, for example), but it makes a nice index of the Dublin Core fields, plus a field for the whole record, for “default” searches. This was a good exercise for me to work with xml, Python and Lucene, because I was having some trouble when trying to index the MODS records for the alternative opac.

    The SRU server is, basically, Dan Chudnov‘s SRU implementation for unalog. It needed to be de-Quixotefied and is, in fact, much more robust than Dan’s original (of course, unalog’s implementation doesn’t need to be as “robust”, since the metadata is much more uniform), but certainly having a working model to modify made this go much, much faster. The nice part is that there might be some stuff in there that Dan might want to put back into unalog.

    So, here is the result. The operations currently supported are explain and searchRetrieve and majority of CQL relations are unsupported, but it does most of the queries I need it to do and, most importantly, it’s less than a week old.

    So the burning question here is: why on earth would I waste time developing this when OCKHAM’s Harvest-to-Query is out there, and, even more specifically, OCLC’s SRW/U implementation for DSpace is available? Further, I knew full well that these projects existed before I started.

    Lemme tell ya.

    Harvest-to-Query looked very promising. I began down this road, but stopped about halfway down the installation document. Granted, anything that uses Perl, TCL and PHP has to be, well, something… After all, those were the first three languages I learned (and in the same order!). Adding in IndexData’s Zebra seemed logical as well since it has a built-in Z39.50 server. Still, this didn’t exactly solve my problem. I’d have to install yazproxy, as well, in order to achieve my SRU requirement. Requiring Perl, TCL, PHP, Zebra and yazproxy is a bit much to maintain for this project. Too many dependencies and I am too easily distracted.

    OCLC’s SRW/U seemed so obvious. It seemed easy. It seemed perfect. Except our DSpace admin couldn’t get it to work. Oh, I inquired. I nagged. I pestered. That still didn’t make it work. I have very limited permissions on the machine that DSpace runs on (and no permissions for Tomcat), so there was little I could do to help. This also solved a specific purpose, but didn’t necessarily address any other OAI providers that we might have.

    So, enter RepoMan. Another wheel that closely resembles all the other wheels out there, but possibly with minor cosmetic changes. Let a thousand wheels be invented.

  • Happy Birthday, you anemic ghoul!

    I heard on NPR this morning that today is Nick Cave’s birthday.

    You can tell from my listening habits that I listen to a lot of Nick Cave, but I had never really thought about him having things like birthdays and whatnot.

    He’s 48. He doesn’t look a day over 375, though.

    So, have a smoke, drink some red wine and lament your wretched existence to your unforgiving maker in celebration.

  • The NISO Riots of Ought-Five

    I am at the NISO OpenURL/Metasearch Workshop, and in almost every way it’s going well. I don’t have a whole lot to say about the workshop content itself. It’s very informative and helpful, but I’ve really got nothing of value to add or comment upon.

    Except one thing.

    Today was “Metasearch” day and just after lunch was a panel presentation of David Lindahl and Jeff Suszczynski of the University of Rochester Libraries and David Walker of Cal State San Marcos Library on “Innovative uses of Metasearch”. And these are two inspirations in the field of metasearch development. Anyone with a metasearch engine (and, sadly, I am not one), please look at these schools’ work in regards to how they are leveraging their metasearch implementations.

    Both of them spoke at some length about what sorts of incredible engineering feats that they had to accomplish to overcome the shit environment that we have found ourselves electronically. During the Q&A session, Peter Noerr, of MuseGlobal (a vendor), basically admonished these three visionaries because vendors are creating tools to do all of the work that they are doing.

    Are you kidding me? They are doing the things that they are doing despite the fact that vendors make it damn near impossible for us to easily provide these services.

    There was a very thick tension that set upon the room at this point. You could almost sense that the librarians, finally driven to the edge by their corporate masters, were ready to rise up, break the chains and slay the dragon that torments them so.

    Almost.

    But then, we’re librarians.

    And the tension passed.

    And then we had a nice break. No chairs were thrown. No blood was let.

    Oh well, maybe next time.