Wednesday, March 10, 2010

Updated Profile: Calais

[Originally published 12/10/09]
The Calais Initiative

Company: Thomson Reuters, Inc.
URL: http://www.opencalais.com/
HQ: New York, USA
Products (Primary): The Calais Initiative
Survey Respondents: Tom Tague, Krista Thomas
Vendor Category: NLP

Employees: 50,000
Revenue: US$13.94 Billion
Calais installed base: 7,000 developers, 2,000,000 pieces of content processed per day

Primary Offering:

The Calais Initiative (Calais) comprises several tools for processing text, but the core product is a Natural Language Processing (NLP) engine. When presented with a body of text, the Calais Web service returns the “named entities” (the categories to which the document’s key terms and concepts are assigned), facts, and events it discovers within the document. The relationships between these items are also identified and embedded in the results. Essentially, the results are the Semantic metadata of the document and can be thought of as the document’s “knowledge content,” which can be published and made available for searching and navigation.

On its own, and applied to one or two small, short documents, this might not seem terribly valuable. But deployed on the Web and made available as a free service, Calais is in a position to process massive amounts of data (text, quantitative, graphic, etc.) and extract their knowledge content. Once this task is complete, this content can be searched individually or combined with other similar content and searched in a larger context. This larger context can be based on other Web content, proprietary Thomson Reuters content, a combination of the two or the context of select data sources that may address a specific area of interest.

Ultimately, Calais’s goal is to be the world’s best tool for extracting the structure of any kind of content, recognizing its type, the concepts that are contained, their relationships, and doing so not just within a single file, but across a span of files that could be as large as the Web itself.

Key Differentiators:

Demand from large organizations, including well established publishers, has grown at an unexpectedly high rate. This has led Thomson Reuters to introduce three contract-based versions of Calais in addition to the original free service:
  • Calais Professional - same as the free service but now backed by an SLA and with higher transaction limits.
  • Calais Professional for Publishers - Calais Professional tailored to meet the needs of large scale publishers and tied to an annual contract.
  • ClearForest On-Premise Solutions - ClearForest is the original name of the technology that makes Calais work. Now that it's available as a stand alone application, enterprises will be able to closely tailor the service to their needs, ensure the privacy of their proprietary content and also have access to what's under the hood for even further customization.
Thomson Reuters is another key differentiator – the fact that Calais is sponsored by a global information giant suggests that this entrant will be with us for a long time. Furthermore, at this time Calais is in the final stages of testing its “infinite scalability” initiative, (e.g., cloud computing) designed to address growth in demand and/or spikes in utilization.

Another distinguishing characteristic is the rate at which the service has been adopted (the fact that it’s free is worth repeating). The net effect has been to discard the original projections for usage because demand has so vastly exceeded expectations. Note that until very recently, dema
nd for Calais has existed almost entirely outside of any Thomson Reuters media property. This state of affairs is changing rapidly, with internal inquiries arriving with greater frequency.

Deploying Calais against the vast, professionally developed and controlled content in the Thomson Reuters empire would be a remarkable step in the company’s evolution. After 150 years as a traditional news wire service and publisher, Thomson Reuters’ content could quickly become something not yet fully defined, but possibly far more powerful and useful than what traditional publishers have offered before.

Six/Twelve Month Plans:

In January ’09, Calais is scheduled to launch Release 4, which will open the door to the world of “Linked Data,” a critical step toward fulfilling the promise of the Semantic Web. Essentially, URIs (Uniform Resource Identifiers) allow for the linking of individual data elements, a concept that goes much further than linking containers like files, pages, documents, or databases as we’re accustomed to on the WWW. The Semantic Web term for each pointer that leads to a datum is “dereferenceable URI”.

Wikipedia does a nice job of explaining references and their consequent dereferencing by using house addresses and houses. In this case, a house address is the reference, or pointer. Using this pointer and finding the actual house is the same as dereferencing the address.

In Calais’s case, after extracting the entities (e.g., people, places, companies, etc.) from your content you could then link to (or retrieve for processing by an application) relevant data on DBpedia, The CIA World Fact Book, Freebase, or a rapidly growing number of other compatible data sources. If you’re a talented content producer, the additional leverage that comes from linking to these “external” data could make your offering substantially more useful and in turn, much more valuable.

Let’s build on the example above, where the entities in an original document have been linked to data residing on DBpedia and The CIA World Fact Book. The idea is that the entities extracted from each source can be linked manually, through search results, or as a result of processing by an application. Simply knowing that these entities have an association can be valuable, but the key is that the URI provides a pointer to the specific data – not the file, not the document, and not the database, but to the actual datum, value, or record that’s stored in one of these containers. There’s no longer a need to call an entire file or database, read it to find what you’re looking for and then put it to use. Instead, you call just what you need – the specific data that matter to you.  

This process is faster (read: cheaper in computer processing terms) and those URIs you’ve amassed can be reused by other people and applications because these pointers are durable and they persist – if the data remain in place, then each datum will keep the same individual URI (again: cheaper, highly reliable, and standardized to ensure universal access and use). It’s simply easier to exchange pointers to specific data (dereferenceable URIs) than it is to exchange potentially huge data files or documents.

Once documents and information assets are connected to the Linked Data cloud, deep connections can be made between the entities, facts and events therein. This can, for instance, enable the resolution of complex queries, such as: “Which company boards of directors include CEOs that have been involved in the sup-prime mortgage meltdown?”.

The diagram (not the datasets) is CC-BY-SA licensed. Email comments to Richard Cyganiak at richard@cyganiak.de

Analysis:

Let’s start with the premise that Thomson Reuters has 150 years of experience creating, managing, and presenting content that people want. Over this period, the company has amassed a body of high quality content that’s possibly the largest in the world. This content will continue to grow, but the advent of the Web has unleashed a torrent of content on a genuinely planetary scale. Since this content is outside Thomson Reuters editorial and/or production controls, the company considers it to be “wild” content. This doesn’t mean it’s bad – some of it’s exceedingly good.

Based on the environmental factors below, Calais puts Thomson Reuters in a position to extend its core competencies to include content it controls as well as wild content because:
  • The fundamental nature of publishing and using content is changing.
  • “World Wild Content” will dwarf the content Thomson Reuters controls.
  • Professionally produced content will continue to merit a premium.
  • The Open Access movement and similar efforts by academics, researchers, and other content authors seeking to retain control of their work will continue and grow.
  • Thomson Reuters has extensive experience in every aspect of the content industry.
  • Flexible integration/interoperation of different types of content may provide powerful added value.
Calais is a free service that stands to significantly benefit people and organizations around the world. The terms of use may vary to allow Calais rights to utilize the content’s metadata or not, but unless you’re a major publisher, this won’t be much of an issue. What matters, at least to Thomson Reuters, is that Calais is a very concrete step toward organizing and integrating the vast span of wild content with its own high quality content. Offering customers your own content combined with the very best of free, Web-based content in an easily searched, highly flexible and exceptionally expansive product is a strong competitive advantage that may ensure another 150 years of operation. This is the strategic thrust of The Calais Initiative.

Thursday, January 7, 2010

Company Profile: IYOUIT

Company Profile: IYOUIT

Company: IYOUIT URL: http://www.iyouit.eu/portal/ HQ: Munich, Germany Products (Primary): IYOUIT Survey Respondents: Matthias Wagner Vendor Category: R&D Project

Employees: -- Revenue: -- Installed base: --

Primary Offering:

The only reason IYOUIT isn’t a runaway global success is because it’s still a research project supported by NTT DOCOMO and the Telematica Instituut.

IYOUIT is a very deliberate effort to explore the use of Semantic Web (SW) technology in mobile environments. IYOUIT integrates a wide range of services such as GPS location, location-based points of interest, picture sharing, local weather, messaging, and more. Some of these data are user generated, while other data are generated automatically, and the application goes even further by connecting to services like Flickr and Twitter. Furthermore, the rich mobile experience is complemented by a Web site (https://www.iyouit.eu/portal/) that displays real time updates from IYOUIT users around the world.

The IYOUIT client is made to run on mobile phones that use the S60 operating system, which means just about any high end phone made by Nokia, LG, or Samsung, along with a few models made by Lenovo and Panasonic. The client is lightweight and its interaction with the network has been tuned to minimize the amount of data passed back and forth. This decision was made deliberately to reduce the impact on subscription plans that charge based on device throughput. Processing demands at the device level have been calibrated to reduce overhead while reasoning, ontology management, and processing-intensive functions occur on the network.

Launched June, 2008, IYOUIT’s user base is still small as these things go – in December, 2008 the project has roughly 1,000 users distributed across 50 countries, with most users concentrated in Europe.

Key Differentiators:

When you’re one of a kind, it’s difficult to contrast with existing products, but some fundamental (and remarkable) qualities include the fact that this application works, it’s available for download right now, it genuinely uses SW technology and it’s made for mobile devices. IYOUIT seamlessly combines the mobile experience with context based enhancements delivered by the network and users can even set “triggers” to be alerted when specified conditions are met, e.g., while you’re at your favorite coffee shop you can be alerted when one of your IYOUIT buddies arrives.

Six/Twelve Month Plans:

As a research project, IYOUIT serves as a learning environment and isn’t necessarily tied to commercial delivery schedules. Nonetheless, the team behind IYOUIT certainly has plans and one that could be discussed is the creation of a developer connection. If this effort succeeds, it’s easy to imagine the creation of more applications and in turn, growth in the user base. In fact, the IYOUIT team is counting on open participation and they’re looking forward to new discoveries.

Analysis:

IYOUIT is much more than an intriguing mobile SW application, so let’s broaden our context (fitting, isn’t it?).

  • While the application is presently geared for relatively high-end phones, all those phones use the S60 operating system originally created by Nokia.
  • Presently, Nokia holds about 40% of the global mobile device market and even if this figure is adjusted to reflect just the higher end of Nokia’s product line, that’s still a lot of phones.
  • Samsung, LG, and others combine to increase the potential user base even further.
  • Nokia has a history of SW research and development that dates back to roughly 1996 and equally, the company has a long history of participating in the open source community.
  • Nokia’s recent SW research seems to focus on the creation of application development tools (http://research.nokia.com/research/projects/), which would play into the promise of IYOUIT very nicely.
  • Nokia’s stated corporate strategy is based on its device business, mobile content, and network infrastructure. Offerings like IYOUIT could be a big win for NTT DOCOMO, Nokia, and just about anyone else who can get involved.
  • NTT DOCOMO is based in Japan and while it’s a cliché at this point, the Asian countries are probably still well ahead of the rest of the world when it comes to developing, deploying, and using mobile technology.

Put these factors together and IYOUIT begins to look like the tip of an iceberg – one that will mean big wins for NTT DOCOMO and other global companies and likely, big wins for innovative startups that create valuable products and services for an environment that’s increasingly ready-made to receive them. Wow!

Monday, January 12, 2009

Company Profile: Nstein

Company Profile: Nstein

Company: Nstein URL: http://www.nstein.com/ HQ: Montreal, Canada Products (Primary): Web Content Management, Text Mining Engine, Digital Asset Management, Picture Management Desk Survey Respondents: David Crouy, Christopher Hill Vendor Category: Vendor

Employees: 200 Revenue: $24M Installed base: 115+

Primary Offering:

For the past two years Nstein has been on a mission to integrate its Text Mining Engine (TME) with the Content Management System (CMS) it gained in its acquisition of Eurocortex. After honing TME for several years prior to acquiring Eurocortex, Nstein discovered that many customers either didn’t know what to do with the resulting metadata or there was no way to use the metadata in existing CMS products. By combining Natural Language Processing (NLP) and CMS, Nstein believed it could pursue significant business opportunities and the company’s client list certainly supports their intuition.

TME is a full-blown NLP product capable of extracting and categorizing the metadata contained within documents, specifically the people or organizations, places, and events that are mentioned in these files. Add-on modules are available that provide document summaries, detect sentiment, search for similar documents, as well as topic clustering. The net effect is that customers can process their documents, categorize and provide search capability within the results, and support their Search Engine Optimization (SEO) efforts.

For publishers, using metadata to create links to relevant content (their own or third party) and a search engine “friendly” Web site is important in capturing incremental revenue. In fact, Nstein actively positions itself as a company that seeks to create additional revenue opportunities for its customers – more on this in a moment. It’s not unusual for many or most publishers to rely on human editors to create categories, relevant links, and manage the posting of these links to the title’s Web site. Obviously, paying people to perform this task can become expensive at a time when the publishing industry is already experiencing a high degree of turmoil.

Solutions like Nstein’s and others can help to reduce the expense of human tagging or even introduce tagging where there’s been no human available to perform this task. Utlimately, costs can only be reduced to zero and businesses rely on revenues and profits for success. Nstein is cognizant of this fact and tries to point its customers in the right direction – for example, once a publisher’s content has been tagged it can be tailored to produce a feed based on a person, place, or thing. For some publishers this can represent a new and very welcome revenue stream. Another example is a common trait of NLP technology, which is the publication of additional content links that are related to the primary article on a given page. Again, some publishers will find the resulting performance an improvement over their current state of affairs.

Key Differentiators:

Aside from being the only CMS platform (or at least one of the very few) to have integrated NLP, and the company’s active focus on creating revenue opportunities for their customers, Nstein’s products include a “Picture Management Desk” designed to manage, tag, and categorize high volumes of images. In fact, a series of virtual desks can be set up to process images depending on the inbound channel.

Six/Twelve Month Plans:

For the moment, Nstein will continue its focus on improving its text mining and associated performance. The company has additional goals to extract even more information where possible and it plans to begin exploring the use of “analytical” queries such “Who was involved in car crashes over the last six months?” or “How many times was John Doe mentioned in articles related to crime?”

Analysis:

Nstein clearly has a head start on most, if not all other vendors in the Content Management System game. Its combination of CMS and NLP is a natural evolution in the management and delivery of Web content and any customer seeking a CMS solution generally would do well to take a close look at Nstein’s products. Far from being an early stage company, Nstein has proven itself during the lull after the dotcom bust, which makes a very positive statement about the fundamental value the company offers. Add to this track record the combination of CMS, NLP, Picture Management, and Digital Asset Management, and a broad range of possibilities become quite clear – both for Nstein and its customers.

Monday, January 5, 2009

Company Profile: Zemanta Ltd.

Company Profile: Zemanta Ltd.

URL: http://www.zemanta.com/ HQ: Ljubljana, Slovenia Products (Primary): Zemanta Web Service Survey Respondents: Andraž Tori Vendor Category: Deployer

Employees: -- Revenue: -- Installed base: --

Primary Offering:

Zemanta has leapt onto the Semantic Web stage by launching its NLP-based service for bloggers and other content producers. The net effect is that for users, it’s like having a co-writer constantly suggesting related articles, links, and images, and then wrapping up with a set of recommended tags designed to increase search engine discovery. The steady stream of suggestions provides an abundant stock of references that can be used to enhance an article, blog post, etc. It’s easy to imagine how anyone would find these helpful, and equally, it’s easy to imagine a productivity boost as well.

Zemanta presently comes in several different flavors including add-ons for Firefox and Internet Explorer, a plug-in for Windows Live Writer (a Microsoft desktop application designed to publish content directly to many popular blogging services), a server-side plug-in for those hosting WordPress, Movable Type, or Drupal implementations and finally, as of December ’08, Zemanta offers fee-based access to its API (http://www.zemanta.com/api/) capabilities for automatic in-text linking, categorization, related news, related images, tagging and linking to other semantic databases. By covering each one of these bases, Zemanta has positioned itself to reach just about anyone who has an interest in blogging, writing, or content creation generally, anywhere in the world.

Recognizing the broad nature of its potentially vast user base, Zemanta’s solution is tuned for casual writers who may have looser writing styles when compared to professional columnists or authors. In fact, Zemanta’s focus on individual content producers may prove to be a key strategic decision – most, if not all, NLP products are geared toward team or institutional environments where there may be larger goals and needs to be served. Becoming a valued tool for individual content producers is a very different strategic aim and Zemanta may have wisely selected a market with few, if any, entrants.

Key Differentiators:

Just trying Zemanta is enough of a differentiator – it’s one thing to have the desire to enhance your content but it’s very different and far better to have ready made suggestions at hand and immediately usable. If you’re a content producer, writer, or just someone looking for helpful suggestions when you write, you may be delighted to have Zemanta’s assistance.

Other differentiators include plans to broaden the delivery of Zemanta’s service through more familiar applications and the ability to harvest thoughts and suggestions from emerging repositories of linked data, but these two points require more time and development prior to their general availability.

A critical factor in setting Zemanta apart is that even the most computer-challenged user can take advantage of this solution. There’s no need for corporate IT personnel to get involved, no approval process, etc. – simply download the add-on to your browser, visit your favorite blogging site and start writing. This kind of simplicity destroys a number of barriers to adoption and Zemanta deserves credit for taking this approach.

Six/Twelve Month Plans:

A commercial version of Zemanta is planned which will be made available on a Software as a Service (SaaS) basis. This version will be targeted toward professional content producers, likely those found in the publishing industry. Once launched, more specialized tools will be soon to follow any offerings to the publishing industry. Zemanta also has plans to apply their solution to more than blogging, although for the time being the company prefers to keep these plans private.

Analysis:

Zemanta is an extremely practical and useful tool that writers of all kinds may find helpful. It only takes a few moments to recognize the potential value of this service, not to mention the sheer helpfulness of having something as tedious as tagging performed automatically.

Returning to company’s market entry point, the decision to pursue individuals initially allows it to acquire recognition, a user base, valuable knowledge from real-world experience, and time to hone its offering to razor sharpness before entering professional/corporate markets. These markets will be pursued primarily in the US and UK, which should certainly keep Zemanta busy for some time to come. Writing may never be quite the same again.

Wednesday, December 3, 2008

Changes At The Semantic Business Blog

If you're a regular reader of Semantic Business you'll recognize that things have been changing - I've added a jobs feed, some of the formatting is slightly off (to my chagrin), and I'm planning to start serving ads in a relatively unobtrusive way. All of this is a precursor to my move over to TypePad (preview here), which looks much better suited to my needs going forward. Please bear with me during this time - this isn't a random decision and I'm aware that links are likely to be broken, things may be lost (hopefully just temporarily), and I may spill my coffee. Or not.
I'm also planning to forward my domain www.davidprovost.com and consolidate my online presence. We'll see how this goes...

Bet On Nokia & S60, Not Google & Android

The excitement of Android's launch has died down a bit so this seems like a good time to cast Google's entry against Nokia. Sounds like an unfair contest right? After all, one is a giant in its industry, a leader in R&D and technical innovation, with a long history of support for the developer/open source community, not to mention a globally recognized brand. I'm referring to Nokia, in case you were wondering. I'm not ignoring Apple and the iPhone, but I don't get the impression they're pressing as hard to shape a play that stretches from fundamental infrastructure up to the end user experience - which I do believe has crossed the minds of people at Google & Nokia.
>
Both companies have money, talent, recognition, and deep commercial relationships. Nokia probably has the lead in governmental relationships due to the highly regulated nature of the telecom industry. Google's got search so solidly nailed it's scary. With Android, Google is moving onto Nokia's turf. In the meantime, Nokia's plans appear to call for an increasing emphasis on Web services and certainly mobile content, as demonstrated by its launch of Ovi and its own music store (maybe Nokia's taking a shot at Apple, after all). 
>
I won't dwell on the increasing sophistication of mobile devices or that they'll become a prominent means of Internet access (or even primary for some people), instead, I'll focus on the evolution we're seeing on the WWW toward the SW. Since the SW is simply an extension of the WWW, it's safe to assume that wherever the Web can be reached, the Semantic Web can be reached as well - it's all due to HTTP after all.
>
Here's where the respective paths of these companies begin to diverge. Google's demonstrated its schizophrenic approach to the SW pretty clearly, as I explored in this post. It seems they're quietly exploring the fringes of the SW, but they're probably expending just as much effort in denying the existence of these activities. I haven't looked through any of Android's documentation for SW references, but since I haven't found any elsewhere, it's possible that there simply aren't any.
>
On the other hand, Nokia's involvement in the Semantic Web goes back to at least to 1997, when Ora Lassila wrote a brief note titled "Introduction to RDF Metadata," likely when he was a visiting fellow at W3C. Ora's gone on to play an influential role within Nokia, promoting the SW all along the way. Others within the company have followed suit so that now there's a raft of internal SW projects here and here (they're not all SW projects, but a good number are). So, given the complexity of mobile environments, the differences in platforms, operating systems, regulatory environments, etc., it seems like SW technology might be well suited to service composition and delivery, provisioning issues generally, and offloading processing requirements from the handset.
>
Oh, did I forget to mention that Nokia acquired Navteq over the summer? Frankly, I don't think there's enough of a difference between Google Maps, Navteq, Yahoo! Maps, or others to really make a difference. But the point is that by purchasing Navteq, Nokia no longer needs to rely on Google and it no longer needs to divulge any information regarding how location based services might be developed and deployed. That's knowledge Google will have to acquire on its own. Since I've already indicated in my earlier post about Google that there's no evidence they're recruiting people with expertise in RDF, OWL, ontologies, SPARQL, graphs, triples, etc., it doesn't look to me like they're going to catch up any time soon. As a matter of fact, now that Google's cutting its workforce , I'm guessing the company's appetite for innovation may slow down, meaning that if they want to catch up in any contest with Nokia, they're going to be sucking serious wind.
>
Diversifying away from a reliance on search makes all the sense in the world to me. But can a mobile operating system succeed on its own without its very own fleet of handsets, infrastructure, content beyond maps, software development in what I believe will be a key technology, and deep experience in global telecom regulation? I don't know, but it's not a bet I would take.

Thursday, November 20, 2008

Company Updates and New Profiles

Since publishing On The Cusp at the end of September, a number of companies have contacted me to schedule briefings and quite frankly, I'm very happy to do so. Of course, now I feel self-conscious because I have a bit of a backlog to attend to...
With that said, I'm planning to provide updates on the companies I've profiled while also profiling new companies (or at least, new to me) in the Semantic Web industry. Additionally, as I become aware of new information or discover something interesting, I'll provide updates accordingly. Along these lines, anyone reading this is welcome to contact me about companies they'd like to see profiled, even if it's their own(!)
In the meantime, you'll be seeing new profiles and updates starting in the very near future.
David

Company Profile: Inform Technologies, Inc.

Inform Technologies, Inc.
URL: http://www.inform.com/
HQ: New York NY, USA
Products (Primary): Inform
Contact: Josh Kirschner
Vendor Category: NLP
Employees: --
Revenue: --
Installed base: --
Primary Offering:
It’s safe to say that Inform has built a successful business by using NLP technology to process its customer’s Web sites and then enrich these pages by linking to relevant content published elsewhere on the Web. One look at Inform’s customer list, which includes names like Ziff Davis, The Economist, Wired, and many more serves as solid evidence that the company has a winning product. Each of these customers has decided that it’s in their best interest to seek out relevant external content and then publish it to supplement the original article on the page. 
Increasingly, publishers seem motivated to engage in this practice in the belief that it reinforces their authoritative standing in the eyes of their audience. As would be expected, customers can control what external content is deemed suitable for publication on their site by creating white lists and black lists, as well as whether or not to keep visitors within the publisher’s “family” of media properties.
In practice, using Inform’s solution is easy enough – Inform hosts the back end where the processing is performed, and when a publisher creates an article its submitted through Inform’s API. The article is processed, the (desired) relevant content is identified and linked, and the results are returned to the customer for final publication. Aside from expanding the content presented to visitors, these results also play an important role in Search Engine Optimization (SEO), with customers reporting page views increasing by 10% to 20% and in some cases as high as 25%.
Key Differentiators:
In the course of its existence, Inform created a well developed, professionally maintained taxonomy. This complements the work of the company’s library scientists, linguists, and ontologists and has had the effect of positioning the company well to pursue specific vertical markets such as health. As a result, Inform is prepared to move into select industries and may well do so from a position of strength, unburdened by the need to play “catch up”.
A very interesting (and fully operational) example of just how far Inform’s solution can be extended is found at NewsDaily.com. Reportedly, this site is operated by a single individual who uses a Reuters feed subscription to form the content kernel for processing by Inform. Clicking through the top-level Reuters content leads to pages the clearly include related articles from a wide range of publishers. Readership or audience numbers weren’t supplied, but the low fixed costs of this business suggest that modest advertising success could yield solid revenues for a one man operation.
Six/Twelve Month Plans:
Aside from its potential entry into specific verticals, Inform is considering opportunities in advertising. While it’s easy to imagine the extraction of primary concepts from an article and then associating relevant advertisements, the actual implementation has challenges which Inform has yet to fully define. Aside from these two possibilities, Inform is deep into execution mode and at this point, a primary goal for the company is to continue building on its track record of success.
Analysis:
Inform has a roster of believers who have contracted for their services. In many ways, the business case is fairly easy to express – in a traditional publishing environment there might be a number of editors who spend part of their time tagging stories for a variety of reasons. Time spent tagging means time taken away from other, higher value activities. Inform’s solution reduces this burden on content editors which translates into time savings, cost savings, and higher productivity (arguably, these are all ways to describe the same business result). These savings, combined with an improved audience experience, create a compelling argument for publishers to take a close look at this technology and how it fits with their overall goals.

Thursday, November 6, 2008

I'm Tired of Google's Shell Game

I don't know anyone that can conclusively say what Google's up to with respect to the Semantic Web, so I'm embarking on a small mission to figure this out - or at least shed a little more light. There are more reasons than I can think of for taking a look at this issue, but for starters:
  • Android vs. Nokia's raft of SW efforts related to mobile environments, and Nokia's stated strategy of relying on three revenue streams deriving from handsets, Web services, and mobile content.
  • Google has a lot of smart people, many of whom are hired out of MIT (probably from the same building where the W3C is headquartered).
  • Unconfirmed reports that upon visiting W3C in Cambridge MA one or two years ago, Eric Schmidt commented that there was lots of "good stuff" going on there (but no research contracts were forthcoming).
  • I just searched on Google's US hiring page and got zero (0) results after using the following nine terms (one at a time with no operators): rdf, owl, sparql, ontology, semantic, uri (although url only got one result), linked, triple, graph.
  • Google's a member of the W3C and since May 1, 2008 nineteen (19) individuals with "@google.com" in their email address have posted on the W3C's public mailing lists (the lists that the working groups use). Some are quite prolific, particularly if they're chairing a working group. Lots of focus on geolocation.
My take is that the search terms have either been scrubbed from the posting results, scrubbed from the job descriptions, or Google just isn't hiring anyone with competencies in the nine areas searched on above. I find this last point unlikely for any forward looking company whose reason for being is the Web itself.
I'll persist in my research and report back - I'll also fill in some links to the points above as well.

Friday, October 24, 2008

Web 3.0 - Highlights From The VC Panel

I sat in on the VC panel at Web 3.0 and took note when they discussed what they look for in the deals they fund. Not a lot of surprises, but I think their points are worth repeating:
  • What problem do you solve? I look at this as a 25-words-or-less plain english statement. I've spoken to a tremendous number of entrepreneurs who struggle to do this and it's not easy. I've created these statements in companies where I've worked and it takes time, thought, and a lot of testing to figure exactly what message gets your point across in a way that people can understand. Obviously, the other benefit is that the more precisely you define the problem, the easier it'll be to solve - hopefully!
  • What's it cost to get a customer? In the late 1990's, I spent a lot of time working with companies like Expedia, E*TRADE, Fidelity, and other household names in e-commerce. Regardless of the industry, each successful e-commerce site came to the realization that their customer acquisition cost was a huge issue and one that needed to be monitored constantly. This point applies equally well to consumer facing sites as well as corporate oriented solution vendors.
  • What do you need to believe in? In other words, if you're starting a social networking site, you need to believe that 1) people are essentially social animals that will usually look for ways to connect with each other; 2) you've identified a niche and created an experience that will actually attract and retain enough of an audience to make the venture pay off; 3) the sun will rise tomorrow morning, or whatever else is essential to success. Being able to state these beliefs helps set context and provide perspective, both of which lend themselves to understanding by a third party, like a VC.
  • What are your dependencies? This is distinct from your beliefs, above - these are the things that have to happen to make your idea work. If you're pursuing a solar power idea, some examples might be 1) silicon based solar panels will continue to fall in price and rise in efficiency; 2) there's a finite supply of oil and over the long term its price will continue to rise; 3) because of construction lead times and political issues, nuclear power won't be a competitor any time soon. And so on - in any case, the point is to identify and clearly understand the things that have to exist for your business to succeed.
  • What's the competitive environment? Well, this one's pretty obvious, assuming you know what business you're really in - see the first point above. Be sure to include substitutes and not just other competitors that happen to look just like you.
  • What's the quality of the team? Putting together a business with your best friends from high school may not be such a great idea, unless they happen to be uncommonly well suited to the tasks they've been assigned. Otherwise, you're far better off reaching out to the best people you can find, which means talking to a lot of people you don't know, getting a glimpse of lot of things you don't know (through their eyes), and hopefully, finding people who are so smart and so talented they leave you thinking you'll have to work hard just to keep up with them. Inevitably, investors need to be convinced that the team in front of them can take their idea and execute on it better than anyone else around.
Whether you're experienced and have dealt with all of these issues or you're new and facing them for the first time, these points are worth bearing in mind as reminders or guides. Now go get 'em!

Monday, October 20, 2008

Thoughts on Jupitermedia Web 3.0

This was a small (~200 attendees) but successful conference. Dan Grigorovici invited me and I'm glad he did - I'd actually geared the release date of my report to be in advance of this conference and I was very nicely surprised at how many people said they found it helpful. Some quick thoughts:
  • I was surprised at how many speakers focused on online advertising or Natural Language Processing (NLP). For awhile I began to think I'd missed something in my registration materials, but I got over any misgivings in time for my panel discussion. Fortunately I had the chance to point out the in the US we have a $13 trillion economy and an audience member volunteered that online advertising accounts for roughly $85 billion. In the remaining $12.9 trillion I'm certain we're leaving many opportunities for Semantic Technology untouched.
  • When I gave a brief talk in Cambridge, MA (MA) two nights earlier, I'd been asked if any corporations were exchanging URIs in the course of conducting business. I'm not aware of any large (Global 1,000) companies engaging in this practice although I can see plenty of reasons to do so. During my panel I ran through a very simple example of Kimberly-Clark (KMB) exchanging URIs with Weyerhaeuser (WY) for production, supply chain, and demand management.
  • In a later session, a collection of West Coast VCs provided a very interesting panel discussion on what they look for in the companies they select for investment - more on this in a separate post.
All in all, Web 3.0 was time and money well spent.

Thursday, October 9, 2008

What a Whirlwind!

Response to my report has been surprisingly strong and my thanks to everyone that's written in. Several companies have contacted me to provide a briefing and I plan to write up those conversations and post them here on my blog. I'm not sure of my own next steps but I plan to offer consulting services in the interim - stay tuned and I'll post details shortly.
And special thanks to Paul Miller for having covered my report - he also kindly invited me to do a podcast which you'll find here.
Next week I'm off to the Web 3.0 conference in Santa Clara where I hope to keep learning more about how companies are using Semantic technologies to achieve their strategic goals. If you plan to attend, I hope you'll introduce yourself and say hello.
David

Tuesday, September 30, 2008

REPORT - On The Cusp: A Global Review of the Semantic Web Industry

I'm delighted to announce the availability of my report titled "On The Cusp: A Global Review of the Semantic Web Industry". In it are profiles of 17 industry participants and analysis that highlights the following:
  • The Semantic Web is a commercially competitive technology.
  • Linked data will be extremely valuable - when it's better understood.
  • Natural Language Processing is an important step in deriving value from "World Wild Content".
To download a copy of this report, click Download Semantic_Web_Industry_Review.pdf (1416.1K)">here.
I'll be part of a panel discussion at Jupitermedia's Web 3.0 conference on Thu Oct 16 and I'll hope to see you there!
In the meantime, my thanks to everyone I spoke with at the companies covered in this report - their participation and involvement made this report possible. Lastly, and far from the least, my thanks to Jim Hendler. For the past several years Jim's been a close advisor and a staunch supporter who's never let me down in a pinch.
David

Sunday, September 14, 2008

Speaking At Jupitermedia's Web3.0 Conference Oct 16-17

I'll be on a panel titled Product Marketing, Key Biz Strategies and the findings in my report will be a key part of my comments. If you're planning to be there I hope you'll say hello!

Tuesday, September 2, 2008

Interviews Are Finished: Final List of Vendors Covered

Fri Sep 5 marked the end of the period I budgeted for interviews and data collection and frankly, I'm delighted with the number of companies that have participated and the thoroughness of their responses. The companies that will be covered in my Sep 30 report are:
In the final report, each of these ventures will be profiled individually to describe their primary offering, what it does, near term plans, and commentary. Several patterns have emerged and I'll discuss these in an industry analysis section that will also be in the report. In the meantime, I'm still on schedule to publish on Sep 30. I'll announce the availability of this [free] report here on my blog and I'll also be sure to provide a link to the .pdf.
My sincere thanks go out to the people and companies that took the time to participate in the interviews and then later, for reviewing the profiles for accuracy - their help and support has been central to this process.
David

Tuesday, August 19, 2008

Update: SemWeb Industry Review, Method and Making Progress

It's time for another progress report - I haven't posted any recent entries to this blog because I've been busy with this project. Even so, since my review of the Semantic Web industry will include a description of my method, I figured I'd get a head start by outlining my approach below. A version of this will appear in the report:
  • Picking the companies is never a perfect process, especially for an initial review. Factors that contributed to being covered by this report include being primarily product based as opposed to consulting services based, some kind of sponsorship at a conference (it's an indicator of success & maturity), and being in "general release" and not beta (a fuzzily enforced criterion). At this point I estimate that 18 companies will be covered.
  • Assemble a list of relevant questions designed to uncover basic business issues, product capabilities, trends, and future plans. These questions were reviewed by a small subset of vendors to ensure fairness, relevance and adequate coverage of key issues.
  • Interviews with key personnel. This is really the heart of the process - the discussion, Q&A, and general exchange involved in covering my questions is what provides the illumination that I'm seeking.
  • Writing company profiles based on the discussions and publicly available information. These are all going to be between one to two pages long and they'll contain a description of the company's primary product, what that product does, plans for the next six to twelve months and some light analysis. There'll be other basic information as well.
  • Submit the profiles to the companies for review. In my view, this is a key step to ensure accuracy. Note that the analysis contained in the body of the report will not be made available for review in this fashion.
  • Summary and analysis, which is where I review the company profiles, develop my findings, and state my observations and comments.
I'm very much in the thick of things right now, but there's always time for trivia:
  • 11 interviews complete, 7 profiles written.
  • The companies span time zones ranging from GMT -7 to GMT +9.
  • 11 are headquartered in North America, 7 are headquartered in Europe, 1 is in Korea.
  • August is global vacation month - no kidding. Why do we even bother with Q3?

Wednesday, August 6, 2008

Semantic Web Industry Review - Progress Report

If you're reading this, you may know that I'm engaged in a Semantic Web industry review based on interviews with the leading industry entrants. Once this phase is complete, I'll reflect upon and analyze my interactions and write up the results. My target for publishing this (free) report is the end of September. Here's where I stand:
Vendors
Aduna - interview being scheduled
Cambridge Semantics - interview complete
Franz - interview being scheduled
Garlik - interview being scheduled
Intellidimension - awaiting initial response
Mondeca - interview complete
Ontoprise - interview being scheduled
Ontos - interview complete
Ontotext Lab - interview being scheduled
OpenLink Software - interview complete
Primal Fusion - special case: pre-launch
Saltlux - interview being scheduled
Sandpiper Software - interview being scheduled
Siderean Software - interview being scheduled
Sindice - special case: pre-launch
Talis - interview being scheduled
Thetus - interview scheduled
TopQuadrant - interview being scheduled
Deployers
Dow Jones Client Solutions - interview scheduled
ITA Software - special case: pre-launch
The Calais Initiative - interview complete
Twine/Radar Networks - interview scheduled
Yahoo!/SearchMonkey - awaiting response

Saturday, August 2, 2008

Note To Self - Virtual:Conceptual as WWW:SW

The World Wide Web is the virtualization layer. All digital assets (databases, files, executables, down to the record level) can connect in a consistent, universal, two-way fashion. 
The Semantic Web is the conceptual layer. The virtualized assets can be combined and linked. Those links can be used by people, machines, other (imagine emailing someone a link to a Web page). It's an important concept, because you as a person are more than what's contained in a credit card database, blog, school transcript, etc. With the SW you can link all those assets together and create a more complete picture of yourself.
In an enterprise, this could be a more complete picture of a customer, chemical compound, financial market, etc.
See the ugly picture with non-standard symbols & terminology (I'll replace it with something nicer - full size here):

Friday, August 1, 2008

Industry Report - Brief Update

I've been in radio silence temporarily as I ramp up the organizational process and start speaking with companies. Response to my invitations has been swift, strong, and enthusiastic.  Nearly all the companies I've contact have agreed to participate in my survey and one or two more have written expressing interest in being included. To be fair, the companies that have chosen not to participate are still in pre-release/stealth mode, but I wanted to make sure they at least had the choice to join in. At this point, I've either scheduled conference calls with each company or doing so is in process, so there's been no hesitation on anyone's part to get started.
The two vendors I've already interviewed have been around for several years and their comments reflect their experience. What struck me the most was their evolution has been based on real-world experience that occurred before the Semantic Web was formalized by the W3C. Each founding team came with hands-on exposure to significant IT infrastructure problems and the recognition that Semantic Web technology is very well suited to solving these problems.
What's also come out of these conversations confirms that the differences will be in the details. Bear in mind I've only interviewed two companies so far and they don't compete with each other. But each one has taken subtle but substantial steps to differentiate the results their customers receive and likewise, the value of these results. Sorry, you'll have to wait for the report for me to clarify that rather vague statement.
More updates to come.

Wednesday, July 23, 2008

Invited Companies - SW Industry Survey

I've sent the invitations and I've been surprised at the speed and extent of the acceptances. Cutting straight to the list of vendors I've invited to participate in my industry survey:
Aduna
Cambridge Semantics
Franz
Garlik
Intellidimension
Mondeca
Ontoprise
Ontos
Ontotext Lab
OpenLink Software
Primal Fusion
Sandpiper Software
Siderean Software
Talis
Thetus
TopQuadrant
And the "deployers" (trying to think of a better term...suggestions are welcome):
BBN
Dow Jones Client Solutions
ITA Software
The Calais Initiative
Twine/Radar Networks
Yahoo!/SearchMonkey
I've created lists for industry reviews before and I've always found that having a few basic guidelines or criteria for inclusion can be very helpful. Here are the ones I've used:
  • The entrant has to be a company or part of a company.
  • Professional services firms (consultants) are not included. I want to focus on vendors or makers of products, Web presences, or Web services.
  • Defense contractors like Northrup Grumman, Boeing, etc. are not included. Those markets are too specialized and restrictions related to classified information arise far too early and often to allow real exposition of the issues.
I've tried to keep this simple and I don't guarantee perfection. As a matter of fact, I've always found that it takes several iterations to get the criteria, categories, and key issues right. This is a starting point where I hope we can all learn together and build on successfully.
Gotta go - I've got to get back to a bunch of people to schedule conference calls. I still owe a couple of entries on this blog related to the process of building shareholder value and separately, using a simple investment portfolio model as an analogy for corporate IT portfolios.