Welcome!

Containers Expo Blog Authors: Karyn Jeffery, Yeshim Deniz, Elizabeth White, Jason Bloomberg, Anders Wallgren

Related Topics: Symbian, Containers Expo Blog, @CloudExpo

Symbian: Article

Exclusive Q&A with Rob Weltman, Director of Grid Services, Yahoo!

Cloud-Based Tools Like Hadoop Are Booming Says Yahoo Exec

Cloud-based tools, including large-scale data-intensive computing as offered by Hadoop, are key to the rise and rise of cloud computing. In this wide-ranging Exclusive Q&A with SYS-CON's Cloud Computing Journal, the Director of Grid Services at Yahoo! - Rob Weltman - explains to Jeremy Geelan, Conference Chair of SYS-CON's 1st International Cloud Computing Conference & Expo held last week in San Jose, CA, how analyzing and learning from ever-growing volumes of business data is essential to continuously refining and improving service offerings.

Cloud Computing Journal: Yahoo! has been the largest contributor to the Hadoop project and uses Hadoop extensively in its Web search and advertising businesses. Can you explain a little of the background to that?
Rob Weltman: Yahoo! Search (and before it Inktomi) was a pioneer in using large clusters of commodity computers to speed up the crawling and indexing of Web sites. While working on the architecture and design of the next generation of Web Search crawling and indexing, we came in touch with Doug Cutting and the open source Lucene project for text indexing/search. Lucene contained a distributed file system with integrated computation using the map-reduce paradigm. It looked very promising and appropriate for many data-intensive applications. Hadoop was then split out as its own project. Yahoo! supported Hadoop in a big way, both in contributing to its development as an open source project and in applying it to solve many large-scale data/computation problems in the company.

Hadoop has matured at an amazingly fast pace. From a 20-node cluster two years ago, to many 2,000-node clusters today; from a somewhat embarrassing terasort (a benchmark) performance to the terasort leader; from a no-access control to user- and group-owned files and directories. There is now a high-level language - Pig - that allows you to express complex operations on data in an intuitive way and have them translated into Hadoop map-reduce jobs.

In 2007, Hadoop at Yahoo! was used primarily for research - analyzing enormous volumes of data to find the best algorithms and parameters for selecting search results or ads to present to users. Now it is also a central component in many production operations, including Web Search, ad serving, and personalization.

Cloud Computing Journal: Are cloud-based tools like Hadoop the most important kinds of tools for the future, do you think?
RW: Being able to add capacity as needed without major software or infrastructure changes is clearly important for many organizations. Sharing resources and dynamically allocating more or less to various functions on demand is highly attractive as companies strive to control costs while the computing needs grow and shift. Analyzing and learning from ever-growing volumes of business data is essential to continuously refining and improving service offerings. The ability to quickly explore new algorithms and put them into production will be a competitive advantage for those with the resources to apply them. All of these speak to the importance of Cloud Computing,

Cloud Computing Journal: How important a role does Java play in the project? Is that because of the need to scale horizontally (and massively)?
RW: Hadoop supports programming and scripting in many languages. Hadoop, itself, is written in Java. The language provides strong support for the central infrastructure needs of system and network programming. There is a large body of experience in developing robust, performance-optimized, scalable platforms in Java.

Java provides portability to many hardware and software environments however Hadoop's horizontal scalability is not a result of the choice of language but rather of a design that is strongly focused on fault-tolerance and distribution.

Cloud Computing Journal: Is the Yahoo! Search Webmap still the world's largest Hadoop production application so far as you are aware? Can you share some size data about Webmap with us?
RW: Yes, as far as I know, the Yahoo! WebMap is the largest Hadoop application in production. It uses 2,000+ computers and is still continuously growing. It produces 300TB of data per run, including 1.2 trillion links.

Cloud Computing Journal: How important are Hadoop clusters to Yahoo! Overall? Do your Web search queries depend on them?
RW: Hadoop isn't directly involved in responding to queries typed in by users, but it is responsible for much of the backend work that produces the indexes used to service those queries. If the Hadoop clusters were down, the quality of search results would quickly degrade as the indexes became stale.

Cloud Computing Journal: Who else besides Yahoo! uses Hadoop to run large distributed computations?
RW: Many of the major Hadoop users are listed at http://wiki.apache.org/hadoop/PoweredBy. Facebook has several hundred nodes in a cluster for backend processing and analysis. Quantcast has several thousand cores in a very large cluster. Many companies, including AOL, A9 (Amazon), and IBM have deployed somewhat smaller clusters. It's likely that almost all of the uses involve large quantities of data.

Cloud Computing Journal: Can Hadoop be run on Amazon EC2?
RW: Absolutely! There is a ready-to-run AMI (virtual machine definition for EC2) for Hadoop. Among many others, Powerset (now owned by Microsoft) runs on EC2.

Cloud Computing Journal: What about Sun's Grid Engine - can it also be run on that?
RW: Yes, Hadoop works with Sun's Grid Engine but you lose the benefit of data locality (putting the computation of each piece of a distributed job near the data needed by that piece).

Cloud Computing Journal: Does the Hadoop team have any kind of a blog or forum?
RW: We have a blog at http://developer.yahoo.net/blogs/hadoop/. The team is also heavily engaged in the user and developer Hadoop mailing lists at hadoop.apache.org.

Cloud Computing Journal: Doug Cutting named it after his child's stuffed elephant. Is there any downside to an Enterprise IT tool having the name of a stuffed elephant?
RW: I did get some ribbing during the election period when I wore my Hadoop Summit t-shirt with the elephant on it, but I was able to clarify Hadoop's open source and non-partisan nature.

Cloud Computing Journal: What else have you and your team developed at Yahoo!, in terms of data-analytics applications for example?
RW: The Grid Computing development team at Yahoo! works on the Hadoop core software, the Pig high-level language, the ZooKeeper distributed coordination service, and the Chukwa monitoring and metric analysis system. In addition, it provides various Hadoop add-ons and tools to e.g. facilitate joining of very large data sets or to understand and improve the performance and efficiency of Hadoop jobs. We provide consulting to application teams that develop large-scale Hadoop programs (often involving feature extraction, modeling, optimization, and index creation) but do not produce them ourselves. 

More Stories By Jeremy Geelan

Jeremy Geelan is Chairman & CEO of the 21st Century Internet Group, Inc. and an Executive Academy Member of the International Academy of Digital Arts & Sciences. Formerly he was President & COO at Cloud Expo, Inc. and Conference Chair of the worldwide Cloud Expo series. He appears regularly at conferences and trade shows, speaking to technology audiences across six continents. You can follow him on twitter: @jg21.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@ThingsExpo Stories
SYS-CON Events announced today that MobiDev will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. MobiDev is a software company that develops and delivers turn-key mobile apps, websites, web services, and complex software systems for startups and enterprises. Since 2009 it has grown from a small group of passionate engineers and business managers to a full-scale mobile software company with over 200 develope...
SoftLayer operates a global cloud infrastructure platform built for Internet scale. With a global footprint of data centers and network points of presence, SoftLayer provides infrastructure as a service to leading-edge customers ranging from Web startups to global enterprises. SoftLayer's modular architecture, full-featured API, and sophisticated automation provide unparalleled performance and control. Its flexible unified platform seamlessly spans physical and virtual devices linked via a world...
SYS-CON Events announced today that Alert Logic, Inc., the leading provider of Security-as-a-Service solutions for the cloud, will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. Alert Logic, Inc., provides Security-as-a-Service for on-premises, cloud, and hybrid infrastructures, delivering deep security insight and continuous protection for customers at a lower cost than traditional security solutions. Ful...
Companies can harness IoT and predictive analytics to sustain business continuity; predict and manage site performance during emergencies; minimize expensive reactive maintenance; and forecast equipment and maintenance budgets and expenditures. Providing cost-effective, uninterrupted service is challenging, particularly for organizations with geographically dispersed operations.
As cloud and storage projections continue to rise, the number of organizations moving to the cloud is escalating and it is clear cloud storage is here to stay. However, is it secure? Data is the lifeblood for government entities, countries, cloud service providers and enterprises alike and losing or exposing that data can have disastrous results. There are new concepts for data storage on the horizon that will deliver secure solutions for storing and moving sensitive data around the world. ...
SYS-CON Events announced today TechTarget has been named “Media Sponsor” of SYS-CON's 18th International Cloud Expo, which will take place on June 7–9, 2016, at the Javits Center in New York City, NY, and the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. TechTarget is the Web’s leading destination for serious technology buyers researching and making enterprise technology decisions. Its extensive global networ...
The IoTs will challenge the status quo of how IT and development organizations operate. Or will it? Certainly the fog layer of IoT requires special insights about data ontology, security and transactional integrity. But the developmental challenges are the same: People, Process and Platform. In his session at @ThingsExpo, Craig Sproule, CEO of Metavine, will demonstrate how to move beyond today's coding paradigm and share the must-have mindsets for removing complexity from the development proc...
SYS-CON Events announced today that Commvault, a global leader in enterprise data protection and information management, has been named “Bronze Sponsor” of SYS-CON's 18th International Cloud Expo, which will take place on June 7–9, 2016, at the Javits Center in New York City, NY, and the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. Commvault is a leading provider of data protection and information management...
SYS-CON Events announced today that MangoApps will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. MangoApps provides modern company intranets and team collaboration software, allowing workers to stay connected and productive from anywhere in the world and from any device. For more information, please visit https://www.mangoapps.com/.
The essence of data analysis involves setting up data pipelines that consist of several operations that are chained together – starting from data collection, data quality checks, data integration, data analysis and data visualization (including the setting up of interaction paths in that visualization). In our opinion, the challenges stem from the technology diversity at each stage of the data pipeline as well as the lack of process around the analysis.
Designing IoT applications is complex, but deploying them in a scalable fashion is even more complex. A scalable, API first IaaS cloud is a good start, but in order to understand the various components specific to deploying IoT applications, one needs to understand the architecture of these applications and figure out how to scale these components independently. In his session at @ThingsExpo, Nara Rajagopalan is CEO of Accelerite, will discuss the fundamental architecture of IoT applications, ...
SYS-CON Events announced today that Tintri Inc., a leading producer of VM-aware storage (VAS) for virtualization and cloud environments, will exhibit at the 18th International CloudExpo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, New York, and the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA.
In his session at 18th Cloud Expo, Bruce Swann, Senior Product Marketing Manager at Adobe, will discuss how the Adobe Marketing Cloud can help marketers embrace opportunities for personalized, relevant and real-time customer engagement across offline (direct mail, point of sale, call center) and digital (email, website, SMS, mobile apps, social networks, connected objects). Bruce Swann has more than 15 years of experience working with digital marketing disciplines like web analytics, social med...
A strange thing is happening along the way to the Internet of Things, namely far too many devices to work with and manage. It has become clear that we'll need much higher efficiency user experiences that can allow us to more easily and scalably work with the thousands of devices that will soon be in each of our lives. Enter the conversational interface revolution, combining bots we can literally talk with, gesture to, and even direct with our thoughts, with embedded artificial intelligence, wh...
SYS-CON Events announced today that EastBanc Technologies will exhibit at SYS-CON's 18th International Cloud Expo®, which will take place on June 7-9, 2016, at the Javits Center in New York City, NY. EastBanc Technologies has been working at the frontier of technology since 1999. Today, the firm provides full-lifecycle software development delivering flexible technology solutions that seamlessly integrate with existing systems – whether on premise or cloud. EastBanc Technologies partners with p...
SYS-CON Events announced today BZ Media LLC has been named “Media Sponsor” of SYS-CON's 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. BZ Media LLC is a high-tech media company that produces technical conferences and expositions, and publishes a magazine, newsletters and websites in the software development, SharePoint, mobile development and Commercial Drone markets.
SYS-CON Events announced today that ContentMX, the marketing technology and services company with a singular mission to increase engagement and drive more conversations for enterprise, channel and SMB technology marketers, has been named “Sponsor & Exhibitor Lounge Sponsor” of SYS-CON's 18th Cloud Expo, which will take place on June 7-9, 2016, at the Javits Center in New York City, New York. “CloudExpo is a great opportunity to start a conversation with new prospects, but what happens after the...
WebRTC is bringing significant change to the communications landscape that will bridge the worlds of web and telephony, making the Internet the new standard for communications. Cloud9 took the road less traveled and used WebRTC to create a downloadable enterprise-grade communications platform that is changing the communication dynamic in the financial sector. In his session at @ThingsExpo, Leo Papadopoulos, CTO of Cloud9, will discuss the importance of WebRTC and how it enables companies to fo...
The IoT is changing the way enterprises conduct business. In his session at @ThingsExpo, Eric Hoffman, Vice President at EastBanc Technologies, discuss how businesses can gain an edge over competitors by empowering consumers to take control through IoT. We'll cite examples such as a Washington, D.C.-based sports club that leveraged IoT and the cloud to develop a comprehensive booking system. He'll also highlight how IoT can revitalize and restore outdated business models, making them profitable...
IoT generates lots of temporal data. But how do you unlock its value? How do you coordinate the diverse moving parts that must come together when developing your IoT product? What are the key challenges addressed by Data as a Service? How does cloud computing underlie and connect the notions of Digital and DevOps What is the impact of the API economy? What is the business imperative for Cognitive Computing? Get all these questions and hundreds more like them answered at the 18th Cloud Expo...