Thursday, October 27, 2011

Cloud Aggregators

I discovered most of these via Gary Price's Best Betas presentation at Internet Librarian. As web apps are exploding and replacing desktop ones, I foresee that services that enable cooperation and aggregation amongst web services will be increasingly in demand. Here's my first taste.

Hojoki

Updated: 3/11/12

I just discovered Hojoki so I thought I would update this post & include it. Hojoki isn't terribly different from the other cloud aggregators here; you can add accounts, from Google Docs to Mendeley (which is a cool bonus for researchers who use that service) to Github, and all updates appear in one place. You can create "Projects" which function as folders and can share folders with people. Updates from any service appear rapidly in Hojoki's equivalent of a timeline, making it a great real-time collaboration tool. Right now, the list of services supported is moderate but interesting. It's particularly cool to see Github & Beanstalk support, meaning that this could make for a better code-collaboration tool than the others on this list. However, lack of FTP/WebDav support makes it more limited than, say, Otixo, which is still what I would recommend as a singular desktop for all your files in disparate cloud applications.

results of a Greplin search

Greplin

Named after the Unix grep command that searches for regular expressions, Greplin is less of an aggregator than a personalized search engine. Not personalized in the way that Google, Bing, and some (but not all...see Blekko and DuckDuckGo, which I believe don't alter results based on personal information) other search engines are these days, but in that you give it access to accounts like Gmail, Facebook, Twitter and it indexes the results. The list of services you can index is fairly large. The user interface is minimalist and slick. Overall, it looks well-done and has the largest chance of making it into my everyday Internet usage of anything on this list.
My only gripe with the service thus far is how it orders search results. First comes Mail, then Events, then People, then Files, and finally Streams (Twitter/Facebook accounts...not much different from People, actually). That's almost precisely the opposite order I'd like to see. When I imagine the utility of something like Greplin, there are two basic use cases: "Damn, what was the cool link I saw somewhere but didn't save?" and "Shoot, where is that document I wrote, in Google Docs or Dropbox?" Neither of those use cases involve my Gmail contacts or Calendar events, yet those are the search results that rise to the top to the detriment of more useful items. I figure Greplin is still young and custom result ordering is probably on the way, so I'm not too concerned. But it does point to perhaps a fundamental misconception of what the service is for.

results of a Otixo search

Otixo

Combines my Dropbox and Google Docs in the same place, a great answer for use case #2 above. This service has, by far, the most limited set of third-party sites it can pull content from, but it might be able to focus in on a single task better because of that. A service that tries to be everything to everyone usually becomes bloated and fails (cough...Facebook...cough). You can also add FTP or WebDav sources, which makes it pretty customizable. I could see this becoming a single source for my scattered web design projects.

Primadesk

This is the only service which I ruled out pretty quickly. It may just be further in beta than the others, but the overall design struck me as clunky and there were signs that the service was very buggy. For one thing, when I went to remove my account, I received an unfiltered error message displaying a Java call stack. I'm nowhere near hacker enough to exploit this but still it sounds an alarm bell in my head when a service that can access my Google, Facebook, and Dropbox accounts gives me a peak at its server-side code. It is one thing to be in beta and another to expose users to error messages. It is even worse when I contact your tech support (whose email address was not easy to locate) and receive no reply.
The list of services you can combine is pretty great, though; around the same size as Greplin. I didn't see any WebDav or FTP support like Otixo, though.

Voyurl

I asked for a beta invite from them and haven't heard back. It sounds more akin to Greplin than Otixo and Primadesk, as in it aims to be a personalized search engine rather than a unified cloud file system. The primary things that intrigue me are the browser plugin (judging from the screenshot it's a Chrome plugin, which is great) and the personal analytics. I'm a data-minded person and I use a lot of social sites, such as Goodreads and Last.fm, simply as repositories of information about myself.

Tuesday, October 4, 2011

Automate Each Week

I always try, at least once a week, to find a cumbersome or lengthy procedure that I do frequently and streamline it somewhat.
To give an example, I often shuffle files between my MacBook laptop and work computer. Sometimes I need to work on files from home, sometimes the OS X interface or a certain program I have just makes things easier.
Now, Dropbox is great for this, but it has its limitations. Leaving security aside, one of the sets of files I work on a lot is a web application for recording library statistics. Since the application requires PHP and MySQL, to run it live on my laptop I need the files to be in my localhost web server directory. So I am constantly copying the latest version from Dropbox, pasting it into my web directory, and then replacing the file that connects to MySQL with a different file (since the MySQL logins on our live site and my localhost are necessarily different). Now, I am not a real programmer, but I know enough of the command line to do this operation via Terminal. So I googled how to write a shell script in OS X and made a libraryStatsTransfer.sh file:


#! /bin/sh
cp -R [Web app's Dropbox location] [Web app's web server location]
cp [The localhost MySQL connection] [The other connection, now in web server directory]



Now, I can run this script and save myself a few seconds and a lot of hideous drag-and-dropping (I am very much a keyboard person, if my post on application launchers didn't already tip you off). Sometimes these little automations don't do much, but other times they're huge and completely transform the way you operate. The first time I installed and configured Quicksilver was the latter.
Computing is an easy example, because computers are all about automation. Any program, at its core, is about automating and simplifying a set of frequently performed commands. But there's no reason to limit oneself to that: cooking, commuting, conversation, etc—all of these have the potential to be streamlined or improved. The real difficulty is in finding something to fix. We are so immersed in our everyday routines that sometimes identifying areas for improvement can be difficult. Then, once you've found something rife with automation potential, thinking up the best way to do it usually isn't hard. Google it, read a book about it, or just meditate for a moment. An answer will present itself.

Tuesday, September 13, 2011

Things Everyone Should Know About Statistics

I recently read How to Lie with Statistics and it solidified the need for this post.


Sample Size Matters

OK Cupid does some cool things with their data. Awhile back, the site published a blog post comparing their homo- and hetero-sexual users that debunked some common, bigoted myths: namely, that homosexuals lust after and seek to convert heterosexuals, and that homosexuals are likely to have had numerous sexual partners. The results were striking: only 0.6% of gay men ever searched for straight matches, and only 0.1% of lesbians did the same. Hetero- and homosexuals had the exact same median number of sexual partners. But what really struck me was a one of the hundreds of comments on this particular post (btw, don't ever read comments) that said something along the lines of “Your study must be wrong, because I've met two gay guys and they both said they had hundreds of sexual partners.” It was hardly the only comment that drew a ridiculous conclusion from a sample size far smaller than OK Cupid's enormous user base. But this is how people think: well, this is my experience, and my experience is representative, so it must be universally true. Needless to say, thousands of varied data points is far superior to any individual's anecdotes. But…


Fully Randomized Sampling is a Fiction

…even OK Cupid's massive sample is innately flawed. Why? Because it is not a cross-section of all gay people. It's all gay people subscribed to an online dating site based in America. So we can expect geographic, national, racial, income, educational, technological, and (most important in this case) relationship-status biases. And I'm not singling out OK Cupid here: I have yet to encounter a research study without some kind of sampling bias. I love Pew's Internet and American Life Project, for instance, but most of their surveys are delivered over the phone, which skews things in a significant way given their topics. People such as myself, who have not owned a landline for almost a decade and have never been listed in a phonebook, are invisible to their methodology. The only studies that come close to true randomization are done in scientific laboratories, where independent variables are carefully curtailed. But such experiments hardly represent life in the wild where causality becomes absurdly complex, and thus are limited in terms of extrapolation. And speaking of causality…


Correlation Does Not Imply Causation

I hesitated a bit before talking about this, because C≠C has become a thoughtless soundbite. I frequently hear it misused in totally irrelevant contexts, and taking the maxim too seriously leads to an insurmountable, Humean skepticism. But it must be said: just because two variables appear to be aligned does not mean there is any causal connection between them. Perhaps the best demonstration of this was produced just recently, with Google's tongue-in-cheek Correlate that finds extremely strong correlations between arbitrary sets of searches, or matches a user-drawn curve to different search terms' popularity over time. Some of these are quite meaningful: Google is famously able to trace flu outbreaks better than the CDC by studying search term occurrence geographically.

The message is: find a healthy level of skepticism in relation to all things statistical or be persistently deceived.

Friday, August 26, 2011

My Ideal Library

A thought experiment. Few people would really want to use this library, I imagine.
  • No printing allowed. Saves librarians hundreds of hours troubleshooting printers (I worked at a library where that was the primary use of reference desk hours), saves the networking gurus innumerable headaches, forces patrons into more reasonable and organized modes of storage (to the clouds!), environmentally-friendly. I hate printers.
  • Only open access e-resources. No proprietary vendor databases, let's see if we can make a legitimate research collection out of DOAJ, Bielefeld, PLoS, arXiv, et al. There are tons of resources out there for free, but can faculty live with that limitation?
  • Hours suited to patrons and not employees. I have this hunch—and I can't confirm it because I don't have access to everyone's circulation, reference, and traffic stats—that being closed on the weekends is more motivated by staff's desire for time off than actual usage at many libraries. Not my commuter college, as it so happens, but at many others. My local public library is closed on Sunday, when few people works, but open at 10am on a Tuesday, when most people are at their jobs. That does not make sense to me.
  • Computer & Internet literacy is taken seriously as a subset of information literacy. Too many librarians complain about teaching students to use a word processor or browser. These things are vital now! This is perhaps the element closest to realization in most libraries: many public libraries already provide computer training. Internet literacy (avoiding phishing, the best services to use, how to modify your browser) is still pretty under-taught.
  • The website would more closely resemble Google's home page than arngren.net. Libraries fulfill lots of roles and it's tough deciding what content deserves space on the home page. But tough choices need to be made because currently we're flooding our users with links, obscuring the useful ones amidst piles of drivel. Tabbed interfaces and aggressive use of analytics to pare down rarely used links are called for.
So am I just a curmudgeon? Or do any of these actually make sense?




Addendum

I can't believe I left this off the list: only open source software on the public computers. We can run Ubuntu with Firefox, Chromium, Libre Office, GIMP, Inkscape, etc. installed. There is practically nothing that cannot be done with free software these days (certainly nothing that MS Office can do) and little reason to force our patrons into Microsoft dependency. MS dominates sheerly by default; the minute a professor accepts .odt for an assignment and tells people that there's a free alternative, the advantage is gone. What's more, with the rise of web apps, everything can be done online in an OS-agnostic environment anyways.

Thursday, August 4, 2011

Painful Procedures

I love First Monday. Along with Code4Lib, it is probably my favorite open access—nay, favorite journal, period. But it also epitomizes one thing that's wrong with the web today: almost everything is unreadable. This post is going to piggyback off my How I Read post about Readability and InstaPaper, which really shouldn't be necessary services but totally are.
In First Monday's instance, here is the process I go through every time I encounter an article I'd like to read in full:
  1. Navigate to the article. In an ideal world, this would be the only step.
  2. The goal is to get this article into InstaPaper, where font size will be increased and dynamic so I can read it on my iPod Touch. But InstaPaper cannot handle frames so I can't work with the typical HTML article. I click "Print version" over in the right-hand column of Open Journal Systems.
  3. Cancel the unwanted print dialog that pops up. Now, normally I would click my InstaPaper bookmarklet and be done, but because this is one of those weird functionless pop-ups my bookmark bar is nowhere to be seen.
  4. Copy the URL, paste into the main browser window.
  5. Cancel the second unwanted print dialog that pops up.
  6. Click the InstaPaper bookmarklet. Now I'm ready to read.
Now, this isn't entirely First Monday's fault, and I don't want to blame Open Journal Systems either. OJS is a great and has aided the spread of open access greatly. But they should know better than to expect people to read their lengthy articles in only one place. No one—well maybe not no one, but very few people read entire articles in one place. PDF would be a substitute except PDFs suck - the have a fixed width so if you have a small screen, like my iPod, you have to zoom in and scroll horizontally to read each line. They're also super annoying for people who don't use Chrome, with it's built-in PDF reader, or who use Windows and thus are forced to download Adobe Reader.
It's a frustrating situation and really very few websites are readable as is. Part of the reason I picked Blogger over Wordpress or another service is that I have more control over page width, font size, and mobile display. But that's a topic for another post.
For a related and better treatment, check out Orbital Content over on A List Apart. Cameron Koczon explores InstaPaper, Readability, and what happens in general when content is freed from its context and starts to revolve around its consumers, not its original site. I think this is a trend that will only continue to increase as more social sites, types of aggregators, and great cross-platform services show up.