2009年3月16日 星期一

Hadoop, a Free Software Program, Finds Uses Beyond Search

Hadoop, a Free Software Program, Finds Uses Beyond Search


Published: March 16, 2009

BURLINGAME, Calif. — In the span of just a couple of years, Hadoop, a free software program named after a toy elephant, has taken over some of the world’s biggest Web sites. It controls the top search engines and determines the ads displayed next to the results. It decides what people see on Yahoo’s homepage and finds long-lost friends on Facebook.

It has achieved this by making it easier and cheaper than ever to analyze and access the unprecedented volumes of data churned out by the Internet. By mapping information spread across thousands of cheap computers and by creating an easier means for writing analytical queries, engineers no longer have to solve a grand computer science challenge every time they want to dig into data. Instead, they simply ask a question.

“It’s a breakthrough,” said Mark Seager, head of advanced computing at the Lawrence Livermore National Laboratory. “I think this type of technology will solve a whole new class of problems and open new services.”

Three top engineers from Google, Yahoo and Facebook, along with a former executive from Oracle, are betting it will. They announced a start-up Monday called Cloudera, based in Burlingame, Calif., that will try to bring Hadoop’s capabilities to industries as far afield as genomics, retailing and finance.

The core concepts behind the software were nurtured at Google.

By 2003, Google found it increasingly difficult to ingest and index the entire Internet on a regular basis. Adding to these woes, Google lacked a relatively easy to use means of analyzing its vast stores of information to figure out the quality of search results and how people behaved across its numerous online services.

To address those issues, a pair of Google engineers invented a technology called MapReduce that, when paired with the intricate file management technology the company uses to index and catalog the Web, solved the problem.

The MapReduce technology makes it possible to break large sets of data into little chunks, spread that information across thousands of computers, ask the computers questions and receive cohesive answers. Google rewrote its entire search index system to take advantage of MapReduce’s ability to analyze all of this information and its ability to keep complex jobs working even when lots of computers die.

MapReduce represented a couple of breakthroughs. The technology has allowed Google’s search software to run faster on cheaper, less-reliable computers, which means lower capital costs. In addition, it makes manipulating the data Google collects so much easier that more engineers can hunt for secrets about how people use the company’s technology instead of worrying about keeping computers up and running.

“It’s a really big hammer,” said Christophe Bisciglia, 28, a former Google engineer and a founder of Cloudera. “When you have a really big hammer, everything becomes a nail.”

The technology opened the possibility of asking a question about Google’s data — like what did all the people search for before they searched for BMW — and it began ascertaining more and more about the relationships between groups of Web sites, pictures and documents. In short, Google got smarter.

The MapReduce technology helps do grunt work, too. For example, it grabs huge quantities of images — like satellite photos — from many sources and assembles that information into one picture. The result is improved versions of products like Google Maps and Google Earth.

Google has kept the inner workings of MapReduce and related file management software a secret, but it did publish papers on some of the underlying techniques. That bit of information was enough for Doug Cutting, who had been working as a software consultant, to create his own version of the technology, called Hadoop. (The name came from his son’s plush toy elephant, which has since been banished to a sock drawer.)

People at Yahoo had read the same papers as Mr. Cutting, and thought they needed to even the playing field with their search and advertising competitor. So Yahoo hired Mr. Cutting and set to work.

“The thinking was if we had a big team of guys, we could really make this rock,” Mr. Cutting said. “Within six months, Hadoop was a critical part of Yahoo and within a year or two it became supercritical.”

A Hadoop-powered analysis also determines what 300 million people a month see. Yahoo tracks peoples’ behavior to gauge what types of stories and other content they like and tries to alter its homepage accordingly. Similar software tries to match ads with certain types of stories. And the better the ad, the more Yahoo can charge for it.

Yahoo is estimated to have spent tens of millions of dollars developing Hadoop, which remains open-source software that anyone can use or modify.

It then began to spread through Silicon Valley and tech companies beyond.

Microsoft became a Hadoop fan when it bought a start-up called Powerset to improve its search system. Historically hostile to open-source software, Microsoft nevertheless altered internal policies to let members of the Powerset team continue developing Hadoop.

“We are realizing that we have real problems to solve that affect businesses, and business intelligence and data analytics is a huge part of that,” said Sam Ramji, the senior director of platform strategy at Microsoft.

Facebook uses it to manage the 40 billion photos it stores. “It’s how Facebook figures out how closely you are linked to every other person,” said Jeff Hammerbacher, a former Facebook engineer and a co-founder of Cloudera.

Eyealike, a start-up, relies on Hadoop for performing facial recognition on photos while Fox Interactive Media mines data with it. Google and I.B.M. have financed a program to teach Hadoop to university students.

Autodesk, a maker of design software, used it to create an online catalog of products like sinks, gutters and toilets to help builders plan projects.. The company looks to make money by tapping Hadoop for analysis on how popular certain items are and selling that detailed information to manufacturers.

These types of applications drew the Cloudera founders toward starting a business around Hadoop.

“What if Google decided to sell the ability to do amazing things with data instead of selling advertising?” Mr. Hammerbacher asked.

Mr. Hammerbacher and Mr. Bisciglia were joined by Amr Awadallah, a former Yahoo engineer, and Michael Olson, the company’s chief executive, who sold a an open-source software company to Oracle in 2006.

The company has just released its own version of Hadoop. The software remains free, but Cloudera hopes to make money selling support and consulting services for the software. It has only a few customers, but it wants to attract biotech, oil and gas, retail and insurance customers to the idea of making more out of their information for less.

The executives point out that things like data copies of the human genome, oil reservoirs and sales data require immense storage systems.

2009年3月12日 星期四

The web, twenty years old

Twenty years of the world wide web

What's the score?

Mar 12th 2009
From The Economist print edition

Science inspired the world wide web. Two decades on, the web has repaid the compliment by changing science


Rex Features A proud father

“INFORMATION Management: A Proposal”. That was the bland title of a document written in March 1989 by a then little-known computer scientist called Tim Berners-Lee who was working at CERN, Europe’s particle physics laboratory, near Geneva. Mr Berners-Lee (pictured) is now, of course, Sir Timothy, and his proposal, modestly dubbed the world wide web, has fulfilled the implications of its name beyond the wildest dreams of anyone involved at the time.

In fact, the web was invented to deal with a specific problem. In the late 1980s, CERN was planning one of the most ambitious scientific projects ever, the Large Hadron Collider, or LHC. (This opened, and then shut down again because of a leak in its cooling system, in September last year.) As the first few lines of the original proposal put it, “Many of the discussions of the future at CERN and the LHC era end with the question—‘Yes, but how will we ever keep track of such a large project?’ This proposal provides an answer to such questions.”

Sir Timothy is now based at the Massachusetts Institute of Technology, where he runs the World Wide Web Consortium, which sets standards for web technology. But on March 13th, he will, if all has gone well, have joined his old colleagues at CERN to celebrate the web’s 20th birthday.

The web of life

The web, as everyone now knows, has found uses far beyond the original one of linking electronic documents about particle physics in laboratories around the world. But amid all the transformations it has wrought, from personal social networks to political campaigning to pornography, it has also transformed, as its inventor hoped it would, the business of doing science itself.

As well as bringing the predictable benefits of allowing journals to be published online and links to be made from one paper to another, it has also, for example, permitted professional scientists to recruit thousands of amateurs to give them a hand. One such project, called GalaxyZoo, used this unpaid labour to classify 1m images of galaxies into various types (spiral, elliptical and irregular). This enterprise, intended to help astronomers understand how galaxies evolve, was so successful that a successor has now been launched, to classify the brightest quarter of a million of them in finer detail. More modestly, those involved in Herbaria@home scrutinise and decipher scanned images of handwritten notes about old plant cuttings stored in British museums. This will allow the tracking of changes in the distribution of species in response to, for instance, climate change.

Another novel scientific application of the web is as an experimental laboratory in its own right. It is allowing social scientists, in particular, to do things that would previously have been impossible.

We recently reported two such projects. One used a peer-to-peer moneylending site to show that a person’s physiognomy is a reliable predictor of his creditworthiness (see article). The other, carried out at the behest of The Economist, confirmed anthropologists’ observations about the sizes of human social networks using data from Facebook (see article). A second investigation of the nature of such networks, which came to similar conclusions, has been produced by Bernardo Huberman of HP Labs, Hewlett-Packard’s research arm in Palo Alto, California. He and his colleagues looked at Twitter, a social networking website that allows people to post short messages to long lists of friends.

At first glance, the networks seemed enormous—the 300,000 Twitterers sampled had 80 friends each, on average (those on Facebook had 120), but some listed up to 1,000. Closer statistical inspection, however, revealed that many of the messages were directed at a few specific friends, revealing—as with Facebook—that an individual’s active social network is far smaller than his “clan”.

Dr Huberman has also helped uncover several laws of web surfing, including the number of times an average person will go from web page to web page on a given site before giving up, and the details of the “winner-takes-all” phenomenon whereby a few sites on a given subject grab most of the attention, and the rest get hardly any.

Scientists have therefore proved resourceful in using the web to further their research. They have, however, tended to lag when it comes to employing the latest web-based social-networking tools to open up scientific discourse and encourage more effective collaboration.

Journalists are now used to having their every article commented on by dozens of readers. Indeed, many bloggers develop and refine their essays on the basis of such input. Yet despite several attempts to encourage a similarly open system of peer review of scientific research published on the web, most researchers still limit such reviews to a few anonymous experts. When Nature, one of the world’s most respected scientific journals, experimented with open peer review in 2006, the results were disappointing. Only 5% of the authors it spoke to agreed to have their article posted for review on the web—and their instinct turned out to be right, as almost half of the papers that were then posted attracted no comments.

Michael Nielsen, an expert on quantum computers who belongs to a new wave of scientist bloggers who want to change this, thinks the reason for this reticence is neither shyness or fear of reprisal, but rather a fundamental lack of incentive.

The unsocial scientist

Scientists publish, in part, because their careers depend on it. They keenly keep track of how many papers they have had accepted, the reputations of the journals they appear in and how many times each article is cited by their peers, as measures of the impact of their research. These numbers can readily be put in a curriculum vitae to impress others.

By contrast, no one yet knows how to measure the impact of a blog post or the sharing of a good idea with another researcher in some collaborative web-based workspace. Dr Nielsen reckons that if similar measurements could be established for the impact of open commentary and open collaboration on the web, such commentary and collaboration would flourish, and science as a whole would benefit. Essentially, this would involve establishing a market for great ideas, just as a site such as eBay does for coveted objects.

How to do that is, at the moment, mysterious. Intellectual credit obeys different rules from the financial sort. But if some keen researcher out there has an idea about how to do it, “Information Management: A Proposal” might be an equally apposite title for his first draft. Who knows, there may even be a knighthood in it.