Your real social network

Foxwhisperer

Photo by Law Keven

Stowe Boyd covered an interesting paper on social networks and concluded "the apparent, superficial social network based on following and followers conceals a deeper, sparser social network". Every current service has an incredibly primitive representation of your relationships. You're either friends with somebody, or you're not. Here's what''s missing:

Strength. There's no way to specify how close you are to somebody else.
Time. Is the friendship long-lasting? Have you talked recently?
Context. What other friends is this friend close to? Which circles do they move in?

It's well known that you can use communication data to answer these questions. This implicit approach is better than trying to get people to enter this information manually because:

Convenience. Nobody wants to spend time doing data entry and house-keeping on their network. Doing it automatically solves that problem.
Reliability. You can objectively measure how many emails somebody has sent you, and how many you've returned to them. This removes the subjective element that creeps in if you're asked to rate the strength of a relationship on an arbitrary scale. It also removes the temptation to exaggerate your closeness to someone influential.

So why hasn't anyone done this? There's massive technical barriers to overcome before you can access large stores of email, and big privacy issues. I'm convinced they can be overcome, and that's what I'm doing with Mailana. If you want to see the sort of detailed social graph I'm talking about, Boulder Twits is using the same backend as my email analysis system.

Why I love long pointless books

Eschercrossing

Photo by Regolare

I recently finished Infinite Jest. It's over a 1000 pages, and has no real arc or resolution, but I enjoyed it immensely. Before that I completed the 12 volumes of A Dance to the Music of Time, another sprawling epic without a conventional plot. Apart from literary masochism, or value for money (I picked up Jest second-hand for 50 cents), why read these monsters?

I realized I'm drawn to them because they feel a lot more real than most other fiction. The characters aren't driven to make decisions that the plot requires. Instead they're set loose on the stage, free to behave randomly, like people. That means you never reach a satisfying conclusion, but then my own life has never had clear-cut resolutions either.

I've also been on a Dickens streak recently, but that's mostly been motivated by the wonderful background characters that weave in and out of the stories. His protagonists and villains are clearly being maneuvered according to the author's plan, making them stilted and artificial. He's not constrained when he's sketching the unimportant people, so they can act like human beings. Little Nell is an alien, but I believed in The Marchioness and Dick Swiveller, despite their lack of purpose.

These books help remind me life's about the journey, not the destination. If you want to make part of your trip more pleasant, I'd recommend picking a long pointless book as a companion.

Get a personal map of your social network

I've upgraded Boulder Twits so that everyone listed has their own personal map, in addition to the graphs showing the whole community. They show who you talk to most on Twitter, organized into groups based on who they talk to. As an example, here's my personal graph. There's some clusters that represent different networks I'm in contact with:

Personalmapboulder

The Boulder folks are mostly in a small, tightly-connected pack on one side. It's almost a mini-version of the full community.

Personalmapapple

Only a few of my Apple colleagues are on Twitter, but they're all pretty interconnected too.

Personalmaplegal

Walter Olson is the founder of the Overlawyered legal blog, and as you can see both me and Jeff Nolan are big fans.

Stay tuned, I'll be using this data to answer some questions like "Who are my friends talking to that I should be following?".

True love and statistics

Mathematicallove

Photo by Keng

I ran an analysis of the most frequent correspondents in the Boulder Twits group, and was very happy to see Gwen Bell and Joel Longtine top of the charts. If you don't know their story, they met through Twitter, and will be getting married soon! It's a wonderful romance, and I was so pleased to see solid mathematical proof of their devotion to each other. My own dear Liz is a statistics major, so I know she'll appreciate it too!

Here's the full top 10, ordered by how many tweets were sent or received by each pair. There's a nice mix of friends and colleagues as well as couples:

  1. gwenbell and jlongtine: 222/291
  2. jennyjenjen and pugofwar: 148/210
  3. abatchelor and bfeld 113/115
  4. neogia and wittytwit: 91/137
  5. micah and technosailor: 82/59
  6. brianlburns and kohlmannj: 65/54
  7. heathercapri and wittytwit: 63/53
  8. micah and w1redone: 49/74
  9. briandewitt and jlongtine: 50/49
  10. ewu and jenn: 65/48

Congratulations to Gwen and Joel, long may their tweeting continue.

The D part of R&D

Mousetrap

Photo by Unloveable

Build a better mousetrap, and the world will beat a path to your door

That's completely wrong. If there's one thing I've learnt over my career, it's that technical excellence is just a small part of a product's success. Distribution is probably the most underrated ingredient, followed by a revenue model, marketing, financing and just plain good timing.

I started off in the UK, working in companies that were packed with insanely smart and resourceful engineers. There's a wonderful tradition over there of celebrating scientists and inventors, everything from the Faraday Christmas Lectures to Dambusters. That creates a big pool of people who can build widgets.

What was missing was the ability to turn a widget into a product. Selling things is a lot less prestigious than inventing them, with all sorts of class overtones of gentlemen scientists and grubby tradesmen mixed in. As a result, most of my companies produced wonderful code, but meager revenues.

Here in the US, I've been able to learn from people versed in the dark arts of actually building a company, not just a piece of software. To be honest it's a lot harder, computers are far more predictable than a gang of primates, but it's also amazing when you step back and see it starting to work. Taking an idea and turning it into something that sustains itself, a living breathing business, that's rewarding as hell.

I'm may not be there yet, but I'm having a blast as I shoot for it.

What analyzing digital communications misses

Closenessdiagram

Greg Berry just posted a very interesting comment, touching on a question I've wrestled with.

"…lots of business and life happens off the internet (hard to believe, I
know), but even within the digital confines, there are so many
different planes of communications to track."

Probably the best example of this is your significant other or business partner. If you're often in the same room as them, you probably won't send them as many emails as a direct report who's in another office. If you rely on communication frequency for measuring closeness, you'll underrate those relationships. So how do you work around this problem?

Design your algorithms around the blindspot. Google's search results are nowhere near as good as a dedicated human researcher could produce, but that doesn't matter. They narrow it down to a couple of dozen sites you can manually check. A few bogus results or dubious rankings don't matter because they can easily be spotted and ignored. The equivalent for tools based on automated relationship analysis is giving users the option to edit the strength of relationships to correct the occasional mistake, and always giving people a chance to eyeball any decision before any action is taken by the system.

Pick the right problem domain. I'm fascinated by applications in the business world because the relationships I needed the most help with are right in the sweet spot for email. I've sketched the graph above to show roughly the communication frequencies I've experienced. For different industries and generations the lines will shift and scale, but between Bob in accounting and your boss there's probably a lot of people you exchange a lot of mails with. Stick to problems related to those folks, and email frequency will be a good approximation to closeness.

Be realistic about the results. I think the Boulder Twits communication map is the best guide to the relationships in the local tech scene, but that's mostly because it's the only one. As Gregg says, different styles of communications heavily affect the results, even if you forget about the channels it's missing. Heavy Twitter users are far more likely to end up in the center of the graph than less prolific twits. Chris Wand is entirely missing because he's not on Twitter, even though he's heavily involved in the community. As we pull in more and more channels we'll be able to produce far better analysis, and do a lot of useful things, but we'll never capture all the fractal richness of relationships within our primate packs.

Javascript, the ginger-haired stepchild of the language family

Redhead

Photo by Gold Sardine

Liz asked me yesterday what language Mailana is written in. It took me a while to think about it, but the list is C (low-level Exchange interfacing), C++ (speed-critical string processing), C# (Outlook plugin), PHP (most of the server architecture), SQL (database querying), Actionscript (Flash components) and Javascript (rich Ajaxesque browser functionality). It got me wondering why the latter gets so little credit, out of all of them it's probably my favorite to use.

I found Douglas Crawford's explanations of why it's the world's most popular, and misunderstood language rang very true, but what really caught my eye were some demos written as pure scripts:

http://www.monstropolis.org/intro8.html
http://www.monstropolis.org/intro1.html
http://www.monstropolis.org/intro3.html
http://www.monstropolis.org/intro7.html
http://www.uselesspickles.com/triangles/demo.html (There's something deeply twisted about rendering 3D triangles using CSS style tricks, but I just can't look away)

I don't know when Javascript will be welcome in polite society, but dismiss it at your peril. It's now everywhere and there's a whole generation of self-taught programmers headed your way who know nothing else.

How I built the Boulder Twits graphs

Clockmechanism

Photo by Pierre J.

I knew I wanted to build a map of how people were connected in the Boulder tech scene. The first step was accessing the raw data, in this case all the Twitter messages from the first 60 local people I'd identified. I already had a system set up to rapidly analyze large numbers of email messages for my Mailana startup. It's modular, with different import components that access mail APIs like Exchange's MAPI/RPC, Gmail's IMAP and Outlook's Object Model, all outputting a stream of messages in standard XML form. Using Twitter's API it was pretty easy to build an importer. The only wrinkle was that I had to search for @someone in the message body, and add that to the recipients field in the XML.

That whirred away for a while pulling in the complete message histories into my database, with indices created keyed on the recipients, as well as lots of other values. Sitting on top of that database I've got a Facebook App-style REST API that let me run queries like "Tell me who sent messages to who within this group of people". Running that on the Twitter messages gave me a list that conceptually looked like this:

Alice to Bob : 10 messages sent, 3 messages received
Alice to Charles: 4 messages sent, 7 received

What I actually wanted was a single number for any relationship, a measure of how strongly Alice and Bob are connected. My choice was the lower of the sent or received counts, so in the above case

Alice to Bob: Strength 3
Alice to Charles: Strength 4

I like this method for mail because it excludes bots like Facebook notification addresses that you never reply to, and penalises other sort of unequal relationships, eg ignoring famous people you might have emailed who ignore you. Not that that ever happens to me of course.

So now I had a list of all the relationships in the community, I needed to display them. I wanted something that could be interacted with inside the browser, so I built a Flash component. I'd never written any Actionscript before, but Mark Shepherd's Springgraph example was a great starting point. After a few days of wrestling with the wonders of flex I had something working.

I then wrote a PHP script that accessed the Mailana API to produce the link information, and the output it in an XML form my component could read in. I based it on the format Daniel Mclaren used for his handy Constellation Roamer plugin, since I'd used that before.

For the Boulder Twits site I didn't want to re-run the query every time to generate the XML. Though it only takes a fraction of a second to create, the system's still pre-alpha so I didn't want a production site depending on it. Instead I saved off several versions and pointed the component directly at the cached XML files. I also didn't want to require every viewer to rerun the force-directed layout, so I let each version arrange itself on my machine, saved the positions and paused the simulation by default. If you want to see the simulation running, try clicking the small play icon in the top left and drag a few people around to see the graph compensate.

I had a lot of fun putting this together. To be honest I was looking for a nice cozy code-womb to crawl into for a couple of weeks after draining my extrovert batteries through Defrag and lots of followup travel and meetings. This was just the ticket, now I'm recharged and looking forward to meeting all the people I've discovered through compiling the list!

How a graph can find missing members

Missingtwits

Social network analysis often devolves into pure eye-candy, but I wanted my graph for more than just a pretty face. I already talked about some of the patterns it reveals, but one of my goals was to uncover all the interesting folks I knew I must be missing off the list. How can a graph help with that?

I'd already analyzed who everyone I had listed talked to, to build the initial network. Next I had to analyze all of their recipients' tweets, to see if they'd replied, and how strong the connection was. By picking the most strongly connected outsiders, and placing them in the same graph, I could see where they fit in the network. In the example, their names are underlined so they stand out.

As you can see, I uncovered quite a few people like z3rr0, w1redone and technosailor who are part of the central group, as well as a lot of other well-connected twits. Once I've weeded out the non-locals, I'll update the list.

This technique is general enough to apply to any group with a partial membership you're trying to complete. For a simple case imagine finding new people to follow by analyzing strangers your friends talk to a lot. Stay tuned for more fun with this.

Will privacy through obscurity work?

Hidingdog

Photo by Angel Shark

Jud's latest post on the generation chasm in attitudes to privacy got me thinking. I'm basing my business on the theory that people will trade privacy for utility in the right conditions. Looking around, everyone from Facebook to Twitter to location-aware services like Brightkite are publicly posting all sorts of personal information. My parents get a neighbor to take the mail in when they're away so burglars won't know the house is empty. Now hundreds of thousands of people tell the world when and where they're on vacation. How come we're not all robbed?

Part of the answer is we're in the honeymoon period for the technology. Remember when every email you got was exciting because it was from a real person, not a robot responder or spam, and you could open attachments without worry? Once services go mainstream, malicious people will abuse them and the media will whip up moral panics.

Another part is the expectation that even though your information is technically public, nobody will bother tracking it down. For example, it's not easy to see conversations between people you don't know on the default Twitter interface. People's attitudes will change once the tools for mining that data improve. It's the same with company email. Everybody knows that their boss or IT admin could be reading their email, but it would be so time-consuming that most people have an expectation of privacy. This is the equivalent of the old Microsoft approach of security through obscurity. Though it has a bad reputation, it worked for a long, long time.

My prediction is we'll keep muddling through as always. There will be backlashes against the complete openness of the current web services as stalkers and spammers attack, but being lost in the crowd isn't a terrible strategy. There will be new social conventions, as we figure out a consensus on what's safe to put online, and new access controls, hopefully based on implicit information like who you've communicated with.