Thursday, July 10, 2008

When I Started a SemWeb Company, Part One

In 2004 I started a Semantic Web company that I named Human Element. The company's technology was based on open source from Simile & DSpace. It didn't work out, but when I folded it after six months, I didn't personally fold because of some very clear cut objectives and timeframes I set for myself. If you've ever thought about starting a venture, or pitching a new line of business within an existing venture (I've done that too, and more successfully), my experience may be helpful.
I knew the problem well enough to get started, and I knew how the technology could solve it. Earlier in my career I spent a number of years as a salesman (and it's unbelievable how helpful that experience is to this day). It may seem corny, but I'm the sort that has to actually write out a cold-call script until I memorize it. Here's a verbatim copy the problem description I used:
  • A community of users with an intense (or high) reliance on data, analysis, and synthesis, combined with an equally strong need and motivation to collaborate between people and/or teams.
And here's a paraphrased version of the solution:
  • DSpace allows end users to publish and share data in any electronic format, such as text, spreadsheets, raw data files, images, etc. The Simile project allows end users to create integrated views of disparate data types. These views allow end users to identify relationships in the data and their behavior that might not otherwise be apparent. 
I pitched my idea into the life sciences industry and dug up all my leads from W3C discussion lists, conference speaker lists, and networking through my friends. The industry was easy to pick because I knew it had the problem I was describing, there was already activity in adopting SW technology, and since a lot of the companies were posting big profits, I thought (mistakenly) it'd be relatively easy to get some cash out of these players. Note that everyone I spoke with was relieved to hear and appreciated my plain English approach.
Milestone #1, speaking in plain English, accomplished.
The technology had already been built. It wasn't plug 'n play ready, but that level of stability & usability wasn't difficult to envision. As a business person, every day I thank my lucky stars for open source and here's why:
  • It's free, which means I don't have to raise investment capital and get locked into commitments to investors. That flexibility is a huge plus. (By the way, my basic approach was BSD=good, GPL=bad, in case you were wondering.)
  • I'll accept the social contract intrinsic to open source any day of the week. Returning improvements to the code in return for the huge head start provided by an existing, free body of code is an incredibly easy decision to make. Any point decisions related to proprietary connectors, implementation tools, interacting with the open source community, etc. were going to be my CTO's job.
  • I'm not technical, so having a body of code already developed eliminates sales objections like "does your product exist anywhere but your own mind?" This factor contributed directly to my 100% track record in getting agreements to demo meetings (when the time came that I actually had a demo.)
Milestone #2, technology that's actually up & running, with a group of people actively maintaining it, accomplished.
I got to know the open source community leaders and enlisted their support. This was pretty important me - I remember when even a hint of commercialism on the Web resulted in flames, excoriation, and eternal damnation, so I felt I had to be able to state unequivocally that I made my intentions known and that I was committed to working with the community. Since the origins of these projects were academic, I spoke to the Principal Investigators, who at the time were David Karger, Eric Miller, and MacKenzie Smith. I got the clear impression they'd have been delighted if I hit a home run with my idea. 
Milestone #3, support and input of the open source leaders, accomplished.
Part Two coming soon, when I'll discuss the money thing and later, recruiting a technical partner, and realizing I had to fold my venture. Stay tuned.

Thursday, July 3, 2008

Is SemWeb on Adoption Wave 1 or 2? Yes.

A close advisor and luminary in SW circles (OK, it's Jim Hendler) described overhearing a discussion at the Semantic Technology Conference '08 where some people were saying that SW technology (SWT) was "...on the first upslope of the "hype curve" and others saying we were on the second (i.e. after the disillusionment)." I saw Jim's story as alluding to the adoption of the SW and chuckled.
Innovation in Practice, Not Theory
There's been a lot written about the process of innovation in the past few decades by people like Clay Christensen, Ed Roberts, Stefan Thomke, Jim Utterback, Eric von Hippel, and many others. One of my favorite examples is Utterback's description of the typewriter, the hurdles it faced when it was introduced, the emergence of a dominant design, and the amount of time it took to really achieve "adoption." You can read Jim's book, but the short version goes a little like this:
  • Initially, typewriters only wrote in ALL CAPS, which offended many recipients because that's the same way handbills (printed promotional messages, or today's equivalent to commercials) were produced.
  • The typist couldn't actually see the output until a couple of lines had been written. This led to error-ridden output.
  • The internal mechanics were slow and unreliable.
  • As different manufacturers entered the market, different keyboard layouts were used.
It took about 30 years to overcome these problems. Innovation takes time, and so does adoption.
Curves, Scales, and Relevant Timeframes
Anyone who's ever written a business plan is always looking for a hockey stick. You know, the kind of curve that starts with a steady, gentle slope, but then explodes upward signifying hyper-adoption and naturally, enormous wealth for all involved parties. 
I think this curve is a pretty good example of what I'm talking about:
OK, I'm being unfair - this is actually the Dow Jones Industrial Average from 1955-2000 (on a linear, not logarithmic scale.) The point I'm trying to make is that all the smooth curves we're accustomed to seeing actually comprise a lot of much smaller curves that are occasionally steep and occasionally precipitous. (I'm also not going to get into a tedious argument about how this might reflect the adoption of financial assets, either.) A lot happened in that 45 year period.
Where Are We Now?
If you're looking to get rich quick, then buy a lottery ticket or attend one of those no-money-down real estate seminars. I don't do either, so I take the long view, which has worked pretty well for Warren Buffet and John Templeton. As a result, I look at the SW industry this way:
In this view, whether we're on curve one or two doesn't matter (and note there's no scale.) The real proposition that we all need to assess is whether or not this is the way we want to invest a significant portion of our time, energy, credibility, and money. My attitude is that it's still very early in the game and I'm quite confident in my bet.

Tuesday, July 1, 2008

Powerset: Breaking the Semantic Web Impasse? What Rubbish.

Microsoft's acquisition of Powerset makes perfect sense to me. After all, MS Word, Excel, PowerPoint, and Outlook play a direct role in the creation of an enormous amount of content, all over the world. Add to that the MS Office suite's XML capabilities and I can begin to imagine some very real and very interesting possibilities when combined with NLP. But is it the Semantic Web? No, but it's definitely a step toward preparing content for use by the Semantic Web, and that's a really good thing.
Natural Language Processing Will Save the Semantic Web!
I took a look at Barney Pell's presentation at ISWC 2007 (but it froze half way through) and while I'm very interested, Natual Language Processing (NLP) simply isn't going to be the panacea he makes it out to be. The description of Barney's talk states that NLP "... can 
break the impasse and open up the possibilities of the Semantic Web." Huh?
Let's get something straight:
  • Natural Language Processing is not the Semantic Web.
  • Semantic Search is not the Semantic Web.
  • Natural Language Processing is a (intriguing) means for structuring written and spoken language so that it can be employed by Semantic Web solutions.
  • There's an enormous amount of data in the world where Natural Language Processing simply won't be applicable.
Don't get me wrong, I like NLP and I like anything that contributes to developing machine readable structures that describe documents. I also like the idea of documents "...as vector-of-keywords" because it seems precisely in line with my intent whenever I create an outline of something I'm going to write. See Barney's slide here:
Complementary Technology or Shared History?
I wouldn't advise Semantic Web startups to start looking at any valuations assigned to Powerset as a guide for the value of their own businesses. With respect to Microsoft, Powerset seems to be a very good fit with the company's portfolio and may play a very interesting role in the evolution of their technology. 
If we really want to join the party, maybe we should harken back to our roots and say it's all artificial intelligence anyway. Doing so might create a big umbrella which would give all our valuations a big boost. But that's admitting that AI is a success, just like expert systems, speech recognition, neural networks, data mining, spam filtering, fraud detection, fuzzy logic...whoops - getting cynical there...

Thursday, June 12, 2008

Don't Hire A Salesperson, Hire Another Developer Instead

Sounds nuts, doesn't it? Note that I'm a firm believer that in a technology company, you either make "it" or you sell "it".
I had lunch with a friend the other day - he's a young guy, he's sharp, and he's launched a Semantic Web venture. He's in that uneasy period where he's got some great technology and a great development team, but he's looking for those first few key customers to start establishing some real traction and also reduce his cash burn.
Let's say you're a founder of a Semantic Web company. If you're resourceful, chances are that identifying leads and starting conversations isn't really the problem. The fact that there are only 24 hours in a day is a much bigger problem. So the question becomes whether or not a salesperson should be brought in to work full time on bringing in customers. For a moment, let's forget that in North America, there might be two to four people who are really qualified to fill this role. What are the options?
  1. Go ahead and hire a full time salesperson, complete with base salary, commission plan, options, benefits, and all the usual stuff. Oh, I forgot, this is an early stage startup and while time is the most critical asset, cash comes in a hairs-breadth later for a close second. That's going to rule out this option pretty quickly unless you're funded by some very deep pockets.
  2. Hire a commission-only, contract salesperson. Be prepared to pay out at a much higher than normal commission rate and don't expect nearly as much control. But those aren't the risks - these are: a) getting this person up to speed on SW technology (good luck); b) risk having that person leave after six months without making a sale (a pro will be ready to move at this point); c) you'll lose six months of precious time and with it... d) the knowledge and relationships that person accrued.
  3. Hire an engineer and learn the sales job yourself. My friend told me he knew for a fact that he's now much better at selling than he was six months ago. I mentioned that someone I really respect once told me that "good managers gravitate to the most difficult (and important) problems" and I really believe that. 
I carried a sales quota for years and there are times when selling can be a tough, unpleasant job. But since I also believe that a company's most senior people are its best sales people, that means if you're a founder, you need to know how to sell if you want to build a business.
If you're a founder, you can't hire a salesperson and expect they'll know the company story anywhere near as well as you do. If you're a founder, it's your job to discover your market and you can't expect to hire anyone to take this responsibility.
Here's the win: Let's say you learn how to sell and find your market. Once you get traction, momentum, whatever you want to call it, there's certain pattern that sets in - you know the questions and answers, the issues and the responses, the competitors, their weaknesses and their strengths. 
You can teach someone else that knowledge, and that's when you hire a salesperson.

Thursday, June 5, 2008

In 23 Words, The Semantic Web Is...Updated Jun 6

The Semantic Web is an extension of the World Wide Web designed to standardize the integration of data and the interoperation of applications.
Update Jun 6: Today I had cause to use this definition and I'm not sure, but my questioner (someone who's very well informed Semantics-wise) seemed slightly aghast at my brevity and simplicity. I understand completely. It's not a glowing definition, full of promise, imbued with excitement, and burgeoning with opportunity. That's a deliberate choice on my part. From what I can tell, fulsome descriptions haven't worked very well over the past few years partly because the technology hasn't been ready and partly because buyers like CIOs, CTOs, and IT managers of all kinds have heard such promises a million times before.
As a business person, I want people outside the Semantic community to buy products made with Semantic technology (the market outside the community is way bigger.) From my sales experience I know that superlatives alone won't cut it and they'll probably just erode your credibility. But what seems very apparent and very real is the technology's ability to standardize the integration of data and the interoperation of applications. Outside of these two things, I'm not aware of anything else Semantically related that's occurring in a full blown, commercial production environment. 
Inferencing/reasoning, SPARQL queries (even with public endpoints), and the flashy stuff that's so promising probably occurs on a daily basis in tightly controlled environments that engage in what I consider "exotic" research and are overseen by the most highly trained and sophisticated people in the world when it comes to this particular technology. That's a far cry from closing a million dollar (euro, yen, yuan, take your pick) sale for an enterprise application that will see widespread deployment to a relatively untrained audience.
When the Semantic Web evolves to the point where it can fulfill all the promises that people (me included) have invested in it, I'll change my definition. But for now, in terms of what I believe it can reliably do today, I'll define the technology as above, although I'll probably still refer to the future capabilities just to keep it interesting.
Besides, I truly believe the emphasis needs to be on the product (e.g. solution) and not its underlying technology. When Microsoft sells Word or Excel, it doesn't emphasize the use of C and Visual Basic (I admit there are exceptions to this) because that's not what people really care about. They want to be able to create documents and work with numbers. It's that simple.