Showing posts with label semantic search. Show all posts
Showing posts with label semantic search. Show all posts

Monday, December 15, 2008

Green Semantic Search: Truevert



The world of semantic search is getting more crowded, and a new entry, Truevert, is taking the concept in a powerful and interesting new direction. Their first offering, Truevert, is a search offering focused exclusively on green issues, and combines semantics with the idea of a baseline interest. It says that rather than finding s search result for an average person without the help of knowing their interest areas leads to results that are irrelevant to many people.

Here's how they describe themselves.

Product overview

The goal of Truevert is to provide users with information that is focused on their interest. Truevert provides a scalable, accurate, and powerful search technology that learns the meaning of language the same way people do - by it's context.

This version of Truevert is focused on green, environmental awareness. All searches are done from the point of view of environmental and social concern. The results are obtained from YAHOO BOSS. They are then organized and clustered by Truevert. If you search for the word "carbon" for example, it knows that you want information about carbon's impact on climate change, not its physical chemistry. What's most surprising, is that the Truevert search engine learned all this in less than one hour on a single server with no ontology, taxonomy, or thesaurus.


For a video intro, see their elevator pitch from Tech Crunch:



Truevert also offers green news and opinion selected for relevance.

I don't know what goes on under the hood of semantic search but this concept makes a lot of horse sense, at least in my eyes. Give it a whirl, it works rather well. I am sure this is just of the first of the semantic vertical search products emanating from here. If your reaction is as positive as mine, they are in for a great future.

Thanks for reading, and don't forget to write.

Thursday, July 31, 2008

Semantic Search Post 1: What is Semantic Search


Whenever I am at a loss for what to post, I try to think about a digital area that I know I should understand better, but have avoided like a plague. For me, semantic search is one of those areas, so I have decided to do ten posts on it. They’re going to flow like this:

1. Intro to Semantic search
2. Company Profile: freebase
3. Company Profile: Powerset
4. Company Profile: Hakia
5. Company Profile: Trovix
6. Responses: Google
7. Responses: Yahoo
8. Conclusions


So let’s get started:

Introduction

Semantic search is a response to a variety of perceived shortcomings in the search space. The idea is that by augmenting search terms with desired meaning, the engine will have a greater likelihood of turning up the best possible results for a consumer.

Semantic search is most relevant to what the Wikipedia entry for it calls “research searches” versus “navigational searches.” In navigational search, the user is looking for a specific object, whereas in research search the user wants to know more about a topic. So for example, if I input “A09275 PDF” I am in all likelihood looking for a copy of the New York bill to regulate BT in that state. If on the other hand, I do a search fro BT privacy, I am likely looking for the best documents to provide insight on the topic. The second type of search is a research search.

Semantic search tools will be of little use in improving the results of navigational searches. When you want something specific, Yahoo and Google offer a fairly high likelihood of pinpointing it online. But for research searches, semantics may prove to help improve search results immeasurably.

The key concept to semantic search is “disambiguation.” By making the goal of a search less ambiguous, a semantic engine could help significantly.

Ambiguity comes in many forms. Consider the phrase “Red China.” While now out of vogue, the term was a popular one 25 or so years ago, and was used as a way of distinguishing Mao’s government in Beijing from Chiang Kai-shek’s in Taipei. But if you typed red china into a search engine you could get sent in a lot of ways:

Red mountains in China
Red dishes
The Communist Government of China
Chinese debt
Etc.


Semantic search helps improves the odds of finding the right document by creating lexical concepts (sorry…meaning “circles.”) and linking different meaning circles to one another based upon meaning similarities.

Such search relies on a relational database (think of a sort of 3-D database versus a flat or 20D database. The 3-D makes more linking and connections possible.

Semantic search also makes it possible to better provide related information. A search for a musician can be followed by a search for lyrics by the musician, etc. But from the marketer’s POV, it simply means more accurate search

More most marketing, semantic search should further improve the effectiveness of online marketing because consumers will be able to more easily find you online. While some have said that semantic search may weaken the graphical ad sector because the web will be less impulse driven through semantic search, I think just the opposite will be the case. The effectiveness of online advertising of all types should increase via this functionality, assuming you count view throughs.

Thanks for reading, and don’t forget to write.

Semantic Search Post 2: Profile: freebase


One of the most interesting and intriguing offerings in the semantic space is freebase, a search platform from Metaweb Technologies. Billing itself as an open database built by a gigantic worldwide of community contributors. Together, members of this community are creating exactly the sort of structured information database that makes better search results available by meaning.

While at first blush, freebase may sound a little like Wikipedia, the crucial difference is in the way that freebase stores information. While Wikipedia is article-based, freebase is, like it says on the tin, a database – a relational database that makes a variety of linkages and conncetions possible through community contributions.

The community contribution aspect is what makes their database different from Google Base, which is organized NOT by the community but rather by Google. It is thus less of a collaboration and more of an expert model, by which I mean that Google essentially rests of the quality of its algoriothms while the efficacy of freebase is directly related to the “wisdom of the crowd.”

Users participate in the community first by browsing the topics available, and by filtering their results according to their specific wants. From there they can add and correct info, or create apps that make the information more useful to searchers.

Finally, users can suggest new schema for storing and organizing the info available. All this makes freebase more dynamic than Wikipedia or Google, though few would doubt the strengths of these offerings as well. Freebase is just a different way to search, use, add to, and deploy info.

Freebase is free to “consumers”, but developers may be charged if they create moneymaking APIs. Ads are also part of the model.

It’s interesting…using freebase takes a little getting used to, because it looks and feels different from Google or Yahoo. But you poke around a bit in it and I bet you will come to the same conclusion that I did – this thing, currently in alpha, could be pretty GD powerful if developer get aboard in a major way.

Thanks for reading, and don’t forget to write.

Semantic Search Post 3: Profile: Powerset

More consumer focused than freebase is in its current form, Powerset is designed to provide quicker and easier access to existing information by topic. Users submit topics or plain English questions and Powerset provides results – currently using Wikipedia and freebase as its underlying databases. You get your results in a classic searchy list, but also simultaneous access to article and information summaries, a “meaning” or “topic” cloud, and other means to quickly zoom in on the facts you want.

The site offers a demo, which I have embedded here:



Powerset Demo Video from officialpowerset on Vimeo.

Essentially, Powerset is a consumer application utilizing the articles in Wikipedia as well as the relational freebase data and surely more such formatted information sources as it goes along. For the average joe, this will be a very useful means of making semantic search relevant and used by a broad set of the population.

Thanks for reading, and don’t forget to write.

Semantic Search Post 4: Profile: Hakia

Hakia offers a different take on semantic search. They describe the key difference between their offering and that of Google or Yahoo by emphasizing quality over popularity of results. Whereas Google uses information popularity (links, etc.) as its surrogate indicator of content quality, Hakia uses a more expert driven approach to identifying the most relevant results. According to their site, quality is defined by Hakia in the following manner:

Quality result satisfies three criteria simultaneously: It (1) comes from credible sources (verticals) recommended by librarians, (2) is the most recent information available, and (3) is absolutely relevant to the query.

Hakia offers a static demo of the difference that quality makes. The example relates to “shifting lanes,” a nautical term that relates in this instance to changing the channels in which ships traverse bodies of water.

I did a few searches of my and found that, as expected, Hakia improves results for research searches. But for navigaitonals, not so much. I found the search results for navigational search on Google far superior. But again, that is as expected. If you want a specific page and are simply using a search engine to navigate there quickly, a regular search platform can be more effective.

Hakia offers site owners two ways to capitalize on semantic search – customized web services and an easy to grab search box.

Additionally, Hakia searches are offered on mobile platforms via Berggi Search, a global mobile search platform.

Thanks for reading, and don’t forget to write.

Semantic Search Post 5: Profile: Trovix


Trovix is an example of putting semantic search to use in a specific vertical, in this case to the job hunt. Here is how the site describes its value to candidates and employers:

For working professionals, we provide a free and effortless way to stay aware of ideal career opportunities. We ensure access to the very best jobs based on work experience and goals, not just keywords. Our service combs through millions of current openings to find and present the most relevant career matches.

Companies use our services to find candidates with the exact skills, experience level, educational background and work history for the positions they want to fill. We then help speed them through the recruiting process with our full-featured applicant tracking system which includes collaborative workflow, streamlined communications, scheduling and custom reporting.

Clearly, their POV is that semantics offer a human-like touch to the issues involved in matching people to jobs and vice versa. Again, from their site:

We take the information contained in a job description as well as a person’s resume and expressed desires, and use that information to rank jobs based on how well they match. With Trovix, it’s like having a personal recruiter who really knows you searching for you and sending only the best jobs or candidates for you.

I like this idea immensely. In fact, I am going to dance a jig to celebrate its arrival. Let me put down the mouse for a sec and do my the dances of my people. Through the miracle of the Internets, I offer the following live webcam footage of me dance. Note: I borrowed the shirt from Jerry Seinfeld.




I’m back. My somewhat silly expression of glee is for the obvious reason. The search results on job sites are mindbogglingly bad. I made a search description using 9 words on one of the major job sites four years ago and never turned it off, in part because I find the dreadful results comical. I used title words, industry words, geography, etc. to describe a desired senior exec job in marketing for a tech company. Today’s best “result” was an offer for a 10-14 hours per week post demoing food in grocery stores.

Trovix analyzes resumes and job descriptions to create tagged versions of each, then matches you to likely relevant results.

Let’s hope the arrival of Trovix revolutionizes the job search sector in a big way.

Thanks for reading, and don’t forget to write.

Semantic Search Post 6: Google Response

Doubtless Google has 927 people working on semantic search, but their first response to this trend, secondary search option, didn’t go down to well with the online retailers that are the bread and butter of their paid search business.

Secondary search works like this: You get a search result and with it a tiny window that lets you search the site of a major retailer as a drill down. Retailers generally ain’t happy because it means that users spend more time on Google and less time on their properties, and because secondary search results can also include competitor ads that could drive people who might otherwise buy from the main retailer to another site.

It’s a very rudimentary semantic approach – it’s essentially applying semantic principles to NAVIGATIONAL search. But there’s gotta be more coming. We’ll see what happened in the months ahead.

This post from TechCrunch points out the tremendous controversy that secondary search has created.

Thanks for reading, and don’t forget to write.

Semantic Search Post 7: Yahoo Response: YOS and Search Monkey

As its core initiatives to improve its search business, Yahoo has taken some dramatic steps to ensure it is out front of the semantic web. Improving search results will be critical because without some major change in the product, Google may well continue to grow its already colossal 60+ share of the market.

YDN is the Yahoo Developer Network, in which companies can use structured search data from yahoo to make apps that will improve the appearance and usefulness of search results. Here’s a vid that explains how developers can use the YDN to make search apps.



Semantic functionality is certainly one of the areas that Yahoo will encourage developers to focus on. Using Search Monkey, the developer can add additional layers of info to search results – including photography, reviews, deep links, etc. Developers will link site owner structured data to consumer search results using microformats or RDF. Developers will be able to add their apps to the yahoo search gallery so consumers can choose the apps that are most appealing to them.

In essence, the quality of search results will be directly connected to the amount and quality of structured data that uis provided by site owners.

Now, while a great deal of the value of YDN and the open platform relates more to navigational versus research searches, the same principles can apply to the research side of searches as well.

While semantics is only a portion of Yahoo’s overall YOS platform, it is a core concept. Additionally, Yahoo will be offering ways for data collected in different Yahoo services (e.g., Yahoo Travel, Fantasy baseball) to enhance individual searches. More relevant results will top the list based upon what consumers have shown interests in in the past.

Thanks for reading, and don’t forget to write.

Semantic Search Post Eight: Conclusions

So search is not my forte, and if I made any errors in all of the other posts, please tell me so I can set the record straight. Additionally, if you hae a company also in this space, please post the url in the comments so I can add it to the list of reviews. I just focused on the outfits that seem to get the most attention.

One of the big conclusions I have about this area is that too many people describe semantic search as a panacea or revolutionary sea change in search. When it appears to me to be an evolutionary addition to existing search principles and information sources. Perhaps I don’t see all the vision though.

What it clear is that search in its current state is astounding, incredible, unbelievably useful, and seriously flawed. We need to remember that only a few years ago search meaqnt a visit to a card catalog, World Book, or microfilm index in the local library. Happily that is no longer the case. But we live in an age of incredible revolutionary and incremental advancements. I view semantic search as the latter, though perhaps it will prove to be both.

Thanks for reading, and don’t forget to write.