Showing posts with label Big data. Show all posts
Showing posts with label Big data. Show all posts

Monday, April 3, 2017

Applied Data Mining and Statistical Learning-Analysis of German Credit Data

source: https://onlinecourses.science.psu.edu/stat857/node/215

Analysis of German Credit Data

Printer-friendly versionPrinter-friendly versiongerman flagData mining is a critical step in knowledge discovery involving theories, methodologies and tools for revealing patterns in data. It is important to understand the rationale behind the methods so that tools and methods have appropriate fit with the data and the objective of pattern recognition. There may be several options for tools available for a data set.
When a bank receives a loan application, based on the applicant’s profile the bank has to make a decision regarding whether to go ahead with the loan approval or not. Two types of risks are associated with the bank’s decision –
  • If the applicant is a good credit risk, i.e. is likely to repay the loan, then not approving the loan to the person results in a loss of business to the bank
  • If the applicant is a bad credit risk, i.e. is not likely to repay the loan, then approving the loan to the person results in a financial loss to the bank

Objective of Analysis:

Minimization of risk and maximization of profit on behalf of the bank.
To minimize loss from the bank’s perspective, the bank needs a decision rule regarding who to give approval of the loan and who not to. An applicant’s demographic and socio-economic profiles are considered by loan managers before a decision is taken regarding his/her loan application.
The German Credit Data contains data on 20 variables and the classification whether an applicant is considered a Good or a Bad credit risk for 1000 loan applicants. Here is a link to the German Credit data (right-click and "save as" ).  A predictive model developed on this data is expected to provide a bank manager guidance for making a decision whether to approve a loan to a prospective applicant based on his/her profiles.
Data Files for this case (right-click and "save as" ) :
The following analytical approaches are taken:
  • Logistic regression: The response is binary (Good credit risk or Bad) and several predictors are available.
  • Discriminant Analysis:
  • Tree-based method and Random Forest
///////////////////////////////////////////////////////////////////////////////////////////////////////

Tuesday, March 14, 2017

Spark 2.0 Technical Preview

source: http://tomining.tistory.com/124

Databricks 에서 게제한 Spark 2.0 Technical Preview 글을 요약해 보았습니다.

spark 1.0 이 공개된 뒤 2년 만에 2.0 release 를 앞두고 있습니다.
Databricks 에서 공개한 Technical Preview 에서는 Spark 2.0의 3가지의 주요 특징을 소개하고 있습니다.

Easier, Faster, Smarter

  • Easier
    • 표준 SQL 지원
      서브쿼리도 지원하는 새로운 Ansi-SQL 파서 적용
    • DataFrame/Dataset API 통합
      • Java/Scala 에서 DataFrame/Dataset 통합
      • SparkSession
        SQLConext 나 HiveContext 를 대체할 DataFrame API 를 위한 진입점
      • 좀 더 간단하고 성능 좋은 Accumlator API
      • 머신러닝 기반의 DataFrame
      • R 을 위한 분산 알고리즘
  • Faster
    • 물리적 실행 영역을 다시 설계
      • CPU 낭비시간 해소
        • 가상함수 호출 시간
        • CPU cache 나 memory 에 데이터를 쓰고 읽는 시간
    • 10억 건을 집계/Join 한 결과


    • Parquet Scan 성능도 3배 이상 개선
  • Smarter
    • Streaming engine 이상의 역할
      외부 저장 시스템(예, RDBMS) 과의 연계, 비즈니스 로직을 잘 처리하는 능력 등
      End-to End “Continuous application” (전체적인 흐름을 아우르는 Application)
    • Structured Streaming API + DataFrame/Dataset API
      실시간 데이터 분석을 가능

아직은 Spark 2.0 이 preview package 이나 몇 주 내로 release 된다고 하니 기대가 됩니다.
Spark Streaming 과 DataFrame/Dataset API 를 잘 활용하면 실시간 분석을 쉽고 간단하게 할 수 있을 것 같습니다.


Wednesday, November 2, 2016

Friday, January 10, 2014

Top 10 Big Data Stories Of 2013

source: http://www.informationweek.com/big-data/big-data-analytics/top-10-big-data-stories-of-2013/d/d-id/1113277?

Top 10 Big Data Stories Of 2013

Big data equals big opportunity -- and a surplus of hype. Catch up on the big data articles that interested readers most in 2013.
 Big Data Talent War: 7 Ways To Win
Big Data Talent War: 7 Ways To Win
(click image for larger view and for slideshow)
Big data ruled as one of the most popular tech topics of 2013, drawing reader interest along many different angles of coverage. Whether focused on careers and education, emerging platforms and technologies, or real-world use cases from healthcare to celebrity social networking, our big data coverage during the last year drew millions of page views.
For a look back at what you may have missed, here's our list of the top-ten big data headlines of 2013.
1. Big Data Analytics Master's Degrees: 20 Top Programs. Our detailed guide to well-known and emerging masters programs specifically targeting the big data analytics talent gap.
2. 7 Big Data Solutions Try To Reshape Healthcare. From Beth Israel Deaconess Medical Center to the University of Pittsburgh Medical Center, leading institutions and technology providers are applying big data to healthcare challenges in innovative ways.
3. 5 Big Wishes For Big Data Deployments. The industry made progress on a couple of these challenges in 2013 -- including SQL-on-Hadoop and stream processing -- but simplified deployment, management, and analysis remain works in progress.
4. Big Data's Surprising Uses: From Lady Gaga To CIA. Lady Gaga's manager created her Littlemonsters.com social network site by mining the singer's 31 million plus fans on Twitter and 51 million plus on Facebook. Now check out the eight other surprising uses covered in this image gallery.
5. Microsoft's Big Data Strategy: An Insider's View. In this in-depth interview, Microsoft executive Dave Campbell outlines plans for Hadoop, machine learning, high-performance computing, and data and analytic offerings on Azure.
6. Big Data Career Switch: 4 Key Points. Looking to retool your skillset to land a job in data science? Beware these issues as you consider university programs.
7. IBM And Big Data Disruption: Insider's View. IBM's Bob Picciano, general manager of Information Management, talks up five big data use cases and Hadoop-driven change -- and  slams SAP Hana and NoSQL databases.
8. NoSQL Vs. Hadoop: Big Data Spotlight At E2. This preview article about a panel discussion at the E2 conference seemed to strike a nerve. The premise: Hadoop is too often seen as a panacea while NoSQL databases are the unsung heroes. Do you agree?
9. Big Data Debate: Will Hadoop Become Dominant Platform? Well-known experts Dave Menninger of Pivotal and James Kobielus of IBM square off on the question of whether Hadoop will become the hub from which most data management activities will either integrate or originate.
10. Big Data: A Practical Definition. Today's hazy definitions don't clearly illustrate big data's benefits. A Hortonworks exec offers a pragmatic alternative.
There's still plenty of debate about just what big data means and whether it will turn out to be an overplayed or underplayed topic where the future of technology is concerned. In our view, data has always been invaluable to effective decision making, and the accuracy of decisions will only improve as we apply more data to important questions. We'll be there to follow the important big data advances in the year ahead. Happy New Year!
Doug Henschen is executive editor of InformationWeek, where he covers the intersection of enterprise applications with information management, business intelligence, big data and analytics. He previously served as editor-in-chief of Intelligent Enterprise, editor-in-chief of Transform Magazine, and executive editor at DM News.
There's no single migration path to the next generation of enterprise communications and collaboration systems and services, and Enterprise Connect delivers what you need to evaluate all the options. Register today and learn about the full range of platforms, services, and applications that comprise modern communications and collaboration systems. Register with code MPIWK and save $200 on the entire event and Tuesday-Thursday conference passes or for a Free Expo pass. It happens in Orlando, Fla., March 17-19.

Thursday, January 9, 2014

Big Data 2014: Powering Up the Curve

source: http://blog.pentaho.com/2013/12/05/quentin-2014-big-data-predictions/

Big Data 2014: Powering Up the Curve

Last year, I predicted that 2013 would be the year big data analytics started to go into mainstream deployment and the research we recently commissioned with Enterprise Management Consultantsindicates that’s happened. What really surprised me though is the extent to which the demand for data blending has powered up the curve and I believe this trend will accelerate big data growth in 2014.
Prediction one: The big data ‘power curve’ in 2014 will be shaped by business users’ demand for data blending
Customers like Andrew Robbins of Paytronix and Andrea Dommers-Nilgen of TravelTainment, who recently spoke about their Pentaho projects at events in NY and London, both come from the business side and are achieving specific goals for their companies by blending big and relational data. Business users like these are getting inspired by the potential to tap into blended data to gain new insights from a 360 degree customer view, including the ability to analyze customer behavior patterns and predict the likelihood that customers will take advantage of targeted offers.
Prediction two: big data needs to play well with others!
Historically, big data projects have largely sat in the IT departments because of the technical skills needed and the growing and bewildering array of technologies that can be combined to build reference architectures. Customers must choose from the various commercial and open source technologies including Hadoop distributions, NoSQL databases, high-speed databases, analytics platforms and many other tools and plug-ins. But they also need to consider existing infrastructure including relational data and data warehouses and how they’ll fit into the picture.
The plus side of all this choice and diversity is that after decades of tyranny and ‘lock-in’ imposed by enterprise software vendors, in 2014, even greater buying power will shift to customers. But there are also challenges. It can be cumbersome to manage this heterogeneous data environment involved with big data analytics. It also means that IT will be looking for Big Data tools to help deploy and manage these complex emerging reference architectures, and to simplify them.  It will be incumbent on the Big Data technology vendors to play well with each other and work towards compatibility. After all, it’s the ability to access and manage information from multiple sources that will add value to big data analytics.
Prediction three: you will see even more rapid innovation from the big data open source community
New open source projects like Hadoop 2.0 and YARN, as the next generation Hadoop resource manager, will make the Hadoop infrastructure more interactive. New open source projects like STORM, a streaming communications protocol, will enable more real-time, on-demand blending of information in the big data ecosystem.
Since we announced the industry’s first native Hadoop connectors in 2010, we’ve been on a mission to make the transition to big data architectures easier and less risky in the context of this expanding ecosystem. In 2013 we made some massive breakthroughs towards this, starting with our most fundamental resource, the adaptive big data layer. This enables IT departments to feel smarter, safer and more confident about their reference architectures and open up big data solutions to people in the business, whether they be data scientists, data analysts, marketing operations analysts or line of business managers.
Prediction four: you can’t prepare for tomorrow with yesterday’s tools
We’re continuing to refine our platform to support the future of analytics. In 2014, we’ll release new functionality, upgrades and plug-ins to make it even easier and faster to move, blend and analyze relational and big data sources. We’re planning to improve the capabilities of the adaptive data layer and make it more secure and easy for customers to manage data flow. On the analytics side, we’re working to simplify data discovery on the fly for all business users and make it easier to find patterns and catch anomalies. In Pentaho Labs, we’ll continue to work with early adopters to cook up new technologies to bring things like predictive, machine data and real-time analytics into mainstream production.
As people in the business continue to see what’s possible with blended big data, I believe we’re going to witness some really exciting breakthroughs and results. I hope you’re as excited as I am about 2014!
Quentin Gallivan, CEO, Pentaho
Big-Data-2014-Predictions-Blog-Graphic

Predicting Big Data's 2014

source: http://www.zdnet.com/predicting-big-datas-2014-7000024189/

Predicting Big Data's 2014

Summary: A summary of Big Data new year predictions from Tableau, Tibco, Alteryx, Basho and Gainsight
There has been no shortage of new year's Big Data and Business Intelligence predictions in my inbox in the past month.  As the predictions have trickled in, I was unsure of the value of running any one slate of predictions as an article in itself.  But once they collected, I realized there's real value in reviewing a compilation of these predictions, seeing where they overlap, and appending a few predictions of my own.
The predictions come from a set of companies in the analytics world ranging from an Enterprise software company (TIBCO, makers of Spotfire); a publicly-traded Business Intelligence company (Tableau); an analytics applications platform start-up (Alteryx); a major NoSQL vendor (Basho); and a customer analytics company (Gainsight).  Together, these companies paint an interesting picture of what the analytics state-of-the-art and market will look like in 2014.
The overarching themes in most of the predictions are: Big Data technologies going mainstream; highly specialized areas of analytics becoming more accessible; an increased influence from cloud and mobile; the continued explosion of data volumes, driven by device- and machine-borne data; and disruption to the incumbent megavendors' hold on the database market.
Mainstream or bust
In the Big Data goes mainstream department, Alteryx predicts 2014 will be when "Hadoop Moves From Curiosity to Critical."  Gainsight imagines a world in 2014 where "People Stop Saying Big Data and Start Meaning It," adding "next year it's table stakes."  TIBCO's CTO, Matt Quinn, says "Big Data and all of its tools and technologies will need to move away from being science experiments, and more into the day-to-day, second-by-second operational decision making."
Another spin on analytics technology becoming more mainstream is that it becomes more applied as well.  Alteryx predicts that in 2014 "Big Data Brings Its 'A Game' in Marketing" and that "Predictive Analytics Will No Longer Be A Specialist Subject." Tableau agrees, surmising that "Predictive Analytics, once the realm of advanced and specialized systems, will move into the mainstream as businesses seek forward-looking rather than backward looking insight from data."
Gainsight says "Businesses Get Proactive by Leveraging Customer Data" and TIBCO's Quinn says "Big Data will focus more on value," explaining that "The challenge is no longer in storing the data...but how to extract value from it."
We're all data scientists nowA couple of companies on our panel are predicting that the high priesthood of data scientists will start to become less crucial as analytics tools become more accessible by business users.  Tableau goes so far as to predict "The end of data scientists," adding "Familiarity with data analysis becomes part of the skill set of ordinary business users, not experts."
Alteryx echos this sentiment with its own "Modern Analysts Matter More Than Data Scientists" prediction, saying "Empowering analysts in business departments with Big Data and analytics will become more important than filling the perceived need for millions of data scientists."  That sounds like a good answer to a problem that I've asserted previously, that "Data scientists don't scale."
Cloud and mobileTableau predicts that "Cloud business intelligence goes mainstream" in 2014, which is a bit ironic since the company only launched its paid cloud solution this year.  It also predicts "Big data finally goes to the sky."  Gainsight says "SaaS Will Become Table Stakes," and although that sentiment refers to the industry as a whole, we can pretty much bet that it applies to the world of Big Data in a prominent fashion.  Basho goes so far as to say "CIOs become cloud operators," expounding that this "will result in many organizations deploying a range of public and private cloud solutions."
TIBCO's view of Big Data and the cloud is encapsulated in its "Big Data will be used for security/cloud" prediction. Quinn expands on that thought, saying "Big Data will be stored and analyzed in the cloud and I expect the bulk of the 2014 data to be generated from cloud services across the stack from PaaS on up."
Moving from cloud to mobile, Tableau says "mobile business intelligence becomes the primary experience, not an occasional experience."  Interestingly, none of the other vendors addressed the question of mobile.  Although Alteryx calls out mobile devices as an important source of data in volume.
Data explosion
The volume and velocity of data shows no sign of letting up, according to most of our predictors.  As I said, Alteryx calls out mobile devices and location data in general as a big source.  Basho has two relevant predictions here, conjecturing that the "'Internet of Things' hits prime time, accelerating the data explosion" and "Your customers' experience will be measured in milliseconds, not seconds."
Measuring customer and market sentiment at such fine granularity has implications for both velocity and volume of data.  Tableau certainly doesn't disagree with Basho's assertion, predicting that "Organizations begin to analyze social data in earnest."
Disruption complexNot surprisingly, Alteryx, a start-up company, reasons that "A new data and analytics stack emerges with new solutions for databases, analytics and visualization, all disrupting the traditional mega-vendors."  Building on the new stack idea, Tableau sees a world in 2014 where "NoSQL technologies become more popular as companies seek ways to assimilate this kind of data."
This sword has two edges though: Basho thinks that the incursion of the new stack into IT will cause IT to demand some help, predicting "Strained Enterprise IT will demand operational simplicity from vendors."  Gainsight might agree, saying 2014 will be a "Big Year for EnterpriseTech IPOs" (emphasis mine).
IMHO, ho, ho
I think the consensus in the predictions we've just gone over serves as a reasonable validation of their efficacy.  I also think our collection of tech companies may have missed a few important potential developments for next year: Hadoop will become more integrated, embedded and less discrete; conventional databases will accommodate NoSQL/semi-structured data workloads, and the industry will consolidate.
With the release of Hadoop 2.0 and its YARN management layer, Hadoop is no longer tied to the MapReduce algorithm and batch processing; as such, it becomes more of a platform and an engine for distributed processing of data, and less of "product."  Additionally, the ability to process key/value data, JSON object/documents, sparse columns and other NoSQL data has already shown up in relational products like DB2 and Vertica, and that trend will continue.
With those opportunities for integration and consolidation of products, I believe it follows that companies will consolidate too.  I expect the Big Data world is teed up for the same kind of M&A activity that the BI world underwent, circa 2007.  
So don't count the enterprise companies out, because they are the likely buyers.  But don't despair, as this is in no way a zero-sum game.  Without all of the innovation of the last few years in the analytics world, the megavendors would likely be giving us another round of relatively minor, evolutionary upgrades in their database platforms.  As things now stand, in 2014, that would be a road that would lead to the industry wilderness.
Topic: Big Data

About 

Andrew J. Brust has worked in the software industry for 25 years as a developer, consultant, entrepreneur and CTO, specializing in application development, databases and business intelligence technology.

Big Data In 2014: 6 Bold Predictions

source: http://www.informationweek.com/big-data/big-data-analytics/big-data-in-2014-6-bold-predictions/d/d-id/1113091



Big Data In 2014: 6 Bold Predictions



'Tis the season when temperatures tumble, shoppers stumble, and prognosticators fumble, often. Will these big data prophecies come true?


How will big data evolve in 2014? The future is anyone's guess, of course, but we thought we'd compile a tasty holiday assortment of prognostications from executives working in the big data trenches. So without further delay, here they are -- six big data predictions for next year:
1. "More Hadoop projects will fail than succeed."That scary assessment is from Gary Nakamura, CEO of Concurrent, a big data application platform company. In a December 12 blog post, Nakamura made a few 2014 forecasts, including this not-so-rosy assessment of Hadoop:


2. Enterprises will focus less on big data and more on stepping up their data management game. 
"There's no doubt that companies' pursuits of big data initiatives have the best intentions to improve operational decision making across the enterprise. That being said, companies shouldn't get stuck on the term 'big data.' The true initiative and what they ultimately need to be concerned with is how they're implementing better data management practices that account for the variety and complexity of the data being acquired for analysis," Scott Schlesinger, a senior vice president for consulting and outsourcing giant Capgemini, told InformationWeek via email.

3. The pace of big data innovation in the open-source community will accelerate in 2014. 
"New open-source projects like Hadoop 2.0 and YARN, as the next-generation Hadoop resource manager, will make the Hadoop infrastructure more interactive. New open-source projects like STORM, a streaming communications protocol, will enable more real-time, on-demand blending of information in the big data ecosystem," wrote Quentin Gallivan, CEO of business analytics software firm Pentaho, in a December 5 blog post.

4. The need for automated tools will become increasingly critical.
"It seems that the more data we have, the more we want," John Joseph, VP of product marketing for analytics software firm Lavastorm Analytics told InformationWeek via email. "But as data volumes increase, the need for pattern matching, simulation, and predictive analytics technologies become more crucial. Engines that can automatically sift through the growing mass of data, identify issues or opportunities, and even take automated action to capitalize on those findings will be a necessity."

5. Beware, Oracle! 2014 will be the year of SQL on Hadoop. 
"I think you'll see people start building interactive applications on the Hadoop infrastructure. And what I mean by that -- and I think this is probably the most controversial thing -- is that people will start replacing their first-generation relational databases with SQL on Hadoop," said Monte Zweben, CEO of SQL-on-Hadoop database startup Splice Machine, in a phone interview with InformationWeek.

6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
2. Enterprises will focus less on big data and more on stepping up their data management game. 
"There's no doubt that companies' pursuits of big data initiatives have the best intentions to improve operational decision making across the enterprise. That being said, companies shouldn't get stuck on the term 'big data.' The true initiative and what they ultimately need to be concerned with is how they're implementing better data management practices that account for the variety and complexity of the data being acquired for analysis," Scott Schlesinger, a senior vice president for consulting and outsourcing giant Capgemini, told InformationWeek via email.
3. The pace of big data innovation in the open-source community will accelerate in 2014. 
"New open-source projects like Hadoop 2.0 and YARN, as the next-generation Hadoop resource manager, will make the Hadoop infrastructure more interactive. New open-source projects like STORM, a streaming communications protocol, will enable more real-time, on-demand blending of information in the big data ecosystem," wrote Quentin Gallivan, CEO of business analytics software firm Pentaho, in a December 5 blog post.

4. The need for automated tools will become increasingly critical.
"It seems that the more data we have, the more we want," John Joseph, VP of product marketing for analytics software firm Lavastorm Analytics told InformationWeek via email. "But as data volumes increase, the need for pattern matching, simulation, and predictive analytics technologies become more crucial. Engines that can automatically sift through the growing mass of data, identify issues or opportunities, and even take automated action to capitalize on those findings will be a necessity."

5. Beware, Oracle! 2014 will be the year of SQL on Hadoop. 
"I think you'll see people start building interactive applications on the Hadoop infrastructure. And what I mean by that -- and I think this is probably the most controversial thing -- is that people will start replacing their first-generation relational databases with SQL on Hadoop," said Monte Zweben, CEO of SQL-on-Hadoop database startup Splice Machine, in a phone interview with InformationWeek.

6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
3. The pace of big data innovation in the open-source community will accelerate in 2014. 
"New open-source projects like Hadoop 2.0 and YARN, as the next-generation Hadoop resource manager, will make the Hadoop infrastructure more interactive. New open-source projects like STORM, a streaming communications protocol, will enable more real-time, on-demand blending of information in the big data ecosystem," wrote Quentin Gallivan, CEO of business analytics software firm Pentaho, in a December 5 blog post.
4. The need for automated tools will become increasingly critical.
"It seems that the more data we have, the more we want," John Joseph, VP of product marketing for analytics software firm Lavastorm Analytics told InformationWeek via email. "But as data volumes increase, the need for pattern matching, simulation, and predictive analytics technologies become more crucial. Engines that can automatically sift through the growing mass of data, identify issues or opportunities, and even take automated action to capitalize on those findings will be a necessity."

5. Beware, Oracle! 2014 will be the year of SQL on Hadoop. 
"I think you'll see people start building interactive applications on the Hadoop infrastructure. And what I mean by that -- and I think this is probably the most controversial thing -- is that people will start replacing their first-generation relational databases with SQL on Hadoop," said Monte Zweben, CEO of SQL-on-Hadoop database startup Splice Machine, in a phone interview with InformationWeek.

6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
4. The need for automated tools will become increasingly critical.
"It seems that the more data we have, the more we want," John Joseph, VP of product marketing for analytics software firm Lavastorm Analytics told InformationWeek via email. "But as data volumes increase, the need for pattern matching, simulation, and predictive analytics technologies become more crucial. Engines that can automatically sift through the growing mass of data, identify issues or opportunities, and even take automated action to capitalize on those findings will be a necessity."
5. Beware, Oracle! 2014 will be the year of SQL on Hadoop. 
"I think you'll see people start building interactive applications on the Hadoop infrastructure. And what I mean by that -- and I think this is probably the most controversial thing -- is that people will start replacing their first-generation relational databases with SQL on Hadoop," said Monte Zweben, CEO of SQL-on-Hadoop database startup Splice Machine, in a phone interview with InformationWeek.

6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
5. Beware, Oracle! 2014 will be the year of SQL on Hadoop. 
"I think you'll see people start building interactive applications on the Hadoop infrastructure. And what I mean by that -- and I think this is probably the most controversial thing -- is that people will start replacing their first-generation relational databases with SQL on Hadoop," said Monte Zweben, CEO of SQL-on-Hadoop database startup Splice Machine, in a phone interview with InformationWeek.
6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
6. Big data flies to the cloud. 
"Big data has gained a lot of traction in 2013 but complex technologies are keeping many businesses from getting their solutions into production and generating a positive ROI. In 2014, businesses will look beyond the hype and turn to cloud solutions that generate fast time to value and do not require highly specialized dedicated skill sets, like Hadoop, to manage. 2014 will be the year that big data moves from buzzword to business imperative," said Sandy Steier, cofounder and CEO of cloud-based analytics firm 1010data, via email.
How will big data evolve in 2014? The future is anyone's guess, of course, but we thought we'd compile a tasty holiday assortment of prognostications from executives working in the big data trenches. So without further delay, here they are -- six big data predictions for next year:
1. "More Hadoop projects will fail than succeed."
"More Hadoop projects will be swept under the rug as businesses devote major resources to their big data projects before doing their due diligence, which results in a costly, disillusioning project failure. We may not hear about most of the failures, of course, but the successes will clearly demonstrate the importance of using the right tools. The right big data toolkit will enable organizations to easily carry forward the success of these projects, as well as further insert their value into their business processes for market advantage."
[Who are next year's rising stars in business intelligence? Read 2014 BI Outlook: Who's Hot, Who's Not.]
Jeff Bertolucci is a technology journalist in Los Angeles who writes mostly for Kiplinger's Personal Finance, the Saturday Evening Post, and InformationWeek.
IT groups need data analytics software that's visual and accessible. Vendors are getting the message. Also in the State Of Analytics issue of InformationWeek: SAP CEO envisions a younger, greener, cloudier company. (Free registration required.)

Monday, September 9, 2013

The Big Data Landscape

source: http://www.forbes.com/sites/davefeinleib/2012/06/19/the-big-data-landscape/

The Big Data Landscape


With the recent IPO of Splunk (currently valued at just over $3 Billion), a lot of attention has turned to Big Data. Problem is, it’s tough to keep track of all the companies involved in the space.
To address that, I’ve created the Big Data Landscape to organize this rapidly growing technology sector. The ecosystem is constantly changing. As such, I welcome your feedback at dave@vcdave.com. I have updated the Big Data Landscape as of July 4, 2012.

Companies, products, and technologies included in the Big Data Landscape:
- Splunk, Loggly, Sumo Logic
- Predictive Policing, BloomReach, Atigeo, Myrrix
- Media Science, Bluefin Labs, CollectiveI, Recorded Future, LuckySort, DataXu, RocketFuel, Turn
- Gnip, Datasift, Space Curve, Factual, Windows Azure Marketplace, LexisNexis, Loqate, Kaggle, Knoema, Inrix
- Oracle Hyperion, SAP BusinessObjects, Microsoft Business Intelligence, IBM Cognos, SAS, MicroStrategy, GoodData, Autonomy, QlikView, Chart.io, Domo, Bime, RJMetrics
- Tableau Software, Palantir, MetaMarkets, Teradata Aster, Visual.ly, KarmaSphere, EMC Greenplum, Platfora, ClearStory Data, Dataspora, Centrifuge, Cirro, Ayata, Alteryx, Datameer, Panopticon, SAS, Tibco, Opera, Metalayer, Pentaho
- HortonWorks, Cloudera, MapR, Vertica, MapR, ParAccel, InfoBright, Kognitio, Calpont, Exasol, Datastax, Informatica
- Couchbase, Teradata, 10gen, Hadapt, Terracotta, MarkLogic, VoltDB,
- Amazon Web Services Elastic MapReduce, Infochimps, Microsoft Windows Azure, Google BigQuery
- Oracle, Microsoft SQL Server, MySQL, PostgreSQL, memsql, Sybase, IBM DB2
- Hadoop, MapReduce, Hbase, Cassandra, Mahout

Big data: What’s your plan?

source: http://www.mckinsey.com/insights/business_technology/big_data_whats_your_plan

Big data: What’s your plan?

Many companies don’t have one. Here’s how to get started.

March 2013| byStefan Biesdorf, David Court, and Paul Willmott
The payoff from joining the big-data and advanced-analytics management revolution is no longer in doubt. The tally of successful case studies continues to build, reinforcing broader research suggesting that when companies inject data and analytics deep into their operations, they can deliver productivity and profit gains that are 5 to 6 percent higher than those of the competition.1 The promised land of new data-driven businesses, greater transparency into how operations actually work, better predictions, and faster testing is alluring indeed.
But that doesn’t make it any easier to get from here to there. The required investment, measured both in money and management commitment, can be large. CIOs stress the need to remake data architectures and applications totally. Outside vendors hawk the power of black-box models to crunch through unstructured data in search of cause-and-effect relationships. Business managers scratch their heads—while insisting that they must know, upfront, the payoff from the spending and from the potentially disruptive organizational changes.
The answer, simply put, is to develop a plan. Literally. It may sound obvious, but in our experience, the missing step for most companies is spending the time required to create a simple plan for how data, analytics, frontline tools, and people come together to create business value. The power of a plan is that it provides a common language allowing senior executives, technology professionals, data scientists, and managers to discuss where the greatest returns will come from and, more important, to select the two or three places to get started.
There’s a compelling parallel here with the management history around strategic planning. Forty years ago, only a few companies developed well-thought-out strategic plans. Some of those pioneers achieved impressive results, and before long a wide range of organizations had harnessed the new planning tools and frameworks emerging at that time. Today, hardly any company sets off without some kind of strategic plan. We believe that most executives will soon see developing a data-and-analytics plan as the essential first step on their journey to harnessing big data.
The essence of a good strategic plan is that it highlights the critical decisions, or trade-offs, a company must make and defines the initiatives it must prioritize: for example, which businesses will get the most capital, whether to emphasize higher margins or faster growth, and which capabilities are needed to ensure strong performance. In these early days of big-data and analytics planning, companies should address analogous issues: choosing the internal and external data they will integrate; selecting, from a long list of potential analytic models and tools, the ones that will best support their business goals; and building the organizational capabilities needed to exploit this potential.
Successfully grappling with these planning trade-offs requires a cross-cutting strategic dialogue at the top of a company to establish investment priorities; to balance speed, cost, and acceptance; and to create the conditions for frontline engagement. A plan that addresses these critical issues is more likely to deliver tangible business results and can be a source of confidence for senior executives.

What’s in a plan?

Any successful plan will focus on three core elements.

Data

A game plan for assembling and integrating data is essential. Companies are buried in information that’s frequently siloed horizontally across business units or vertically by function. Critical data may reside in legacy IT systems that have taken hold in areas such as customer service, pricing, and supply chains. Complicating matters is a new twist: critical information often resides outside companies, in unstructured forms such as social-network conversations.
Making this information a useful and long-lived asset will often require a large investment in new data capabilities. Plans may highlight a need for the massive reorganization of data architectures over time: sifting through tangled repositories (separating transactions from analytical reports), creating unambiguous golden-source data,2 and implementing data-governance standards that systematically maintain accuracy. In the short term, a lighter solution may be possible for some companies: outsourcing the problem to data specialists who use cloud-based software to unify enough data to attack initial analytics opportunities.

Analytic models

Integrating data alone does not generate value. Advanced analytic models are needed to enable data-driven optimization (for example, of employee schedules or shipping networks) or predictions (for instance, about flight delays or what customers will want or do given their buying histories or Web-site behavior). A plan must identify where models will create additional business value, who will need to use them, and how to avoid inconsistencies and unnecessary proliferation as models are scaled up across the enterprise.
As with fresh data sources, companies eventually will want to link these models together to solve broader optimization problems across functions and business units. Indeed, the plan may require analytics “factories” to assemble a range of models from the growing list of variables and then to implement systems that keep track of both. And even though models can be dazzlingly robust, it’s important to resist the temptation of analytic perfection: too many variables will create complexity while making the models harder to apply and maintain.

Tools

The output of modeling may be strikingly rich, but it’s valuable only if managers and, in many cases, frontline employees understand and use it. Output that’s too complex can be overwhelming or even mistrusted. What’s needed are intuitive tools that integrate data into day-to-day processes and translate modeling outputs into tangible business actions: for instance, a clear interface for scheduling employees, fine-grained cross-selling suggestions for call-center agents, or a way for marketing managers to make real-time decisions on discounts. Many companies fail to complete this step in their thinking and planning—only to find that managers and operational employees do not use the new models, whose effectiveness predictably falls.
There’s also a critical enabler needed to animate the push toward data, models, and tools: organizational capabilities. Much as some strategic plans fail to deliver because organizations lack the skills to implement them, so too big-data plans can disappoint when organizations lack the right people and capabilities. Companies need a road map for assembling a talent pool of the right size and mix. And the best plans will go further, outlining how the organization can nurture data scientists, analytic modelers, and frontline staff who will thrive (and strive for better business outcomes) in the new data- and tool-rich environment.
By assembling these building blocks, companies can formulate an integrated big-data plan similar to what’s summarized in the exhibit. Of course, the details of plans—analytic approaches, decision-support tools, and sources of business value—will vary by industry. However, it’s important to note an important structural similarity across industries: most companies will need to plan for major data-integration campaigns. The reason is that many of the highest-value models and tools (such as those shown on the right of the exhibit) increasingly will be built using an extraordinary range of data sources (such as all or most of those shown on the left). Typically, these sources will include internal data from customers (or patients), transactions, and operations, as well as external information from partners along the value chain and Web sites—plus, going forward, from sensors embedded in physical objects.

Exhibit

A successful data plan will focus on three core elements.
To build a model that optimizes treatment and hospitalization regimes, a company in the health-care industry might need to integrate a wide range of patient and demographic information, data on drug efficacy, input from medical devices, and cost data from hospitals. A transportation company might combine real-time pricing information, GPS and weather data, and measures of employee labor productivity to predict which shipping routes, vessels, and cargo mixes will yield the greatest returns.

Three key planning challenges

Every plan will need to address some common challenges. In our experience, they require attention from the senior corporate leadership and are likely to sound familiar: establishing investment priorities, balancing speed and cost, and ensuring acceptance by the front line. All of these are part and parcel of many strategic plans, too. But there are important differences in plans for big data and advanced analytics.

1. Matching investment priorities with business strategy

As companies develop their big-data plans, a common dilemma is how to integrate their “stovepipes” of data across, say, transactions, operations, and customer interactions. Integrating all of this information can provide powerful insights, but the cost of a new data architecture and of developing the many possible models and tools can be immense—and that calls for choices. Planners at one low-cost, high-volume retailer opted for models using store-sales data to predict inventory and labor costs to keep prices low. By contrast, a high-end, high-service retailer selected models requiring bigger investments and aggregated customer data to expand loyalty programs, nudge customers to higher-margin products, and tailor services to them.
That, in a microcosm, is the investment-prioritization challenge: both approaches sound smart and were, in fact, well-suited to the business needs of the companies in question. It’s easy to imagine these alternatives catching the eye of other retailers. In a world of scarce resources, how to choose between these (or other) possibilities?
There’s no substitute for serious engagement by the senior team in establishing such priorities. At one consumer-goods company, the CIO has created heat maps of potential sources of value creation across a range of investments throughout the company’s full business system—in big data, modeling, training, and more. The map gives senior leaders a solid fact base that informs debate and supports smart trade-offs. The result of these discussions isn’t a full plan but is certainly a promising start on one.
Or consider how a large bank formed a team consisting of the CIO, the CMO, and business-unit heads to solve a marketing problem. Bankers were dissatisfied with the results of direct-marketing campaigns—costs were running high, and the uptake of the new offerings was disappointing. The heart of the problem, the bankers discovered, was a siloed marketing approach. Individual business units were sending multiple offers across the bank’s entire base of customers, regardless of their financial profile or preferences. Those more likely to need investment services were getting offers on a range of deposit products, and vice versa.
The senior team decided that solving the problem would require pooling data in a cross-enterprise warehouse with data on income levels, product histories, risk profiles, and more. This central database allows the bank to optimize its marketing campaigns by targeting individuals with products and services they are more likely to want, thus raising the hit rate and profitability of the campaigns. A robust planning process often is needed to highlight investment opportunities like these and to stimulate the top-management engagement they deserve given their magnitude.

2. Balancing speed, cost, and acceptance

A natural impulse for executives who “own” a company’s data and analytics strategy is to shift rapidly into action mode. Once some investment priorities are established, it’s not hard to find software and analytics vendors who have developed applications and algorithmic models to address them. These packages (covering pricing, inventory management, labor scheduling, and more) can be cost-effective and easier and faster to install than internally built, tailored models. But they often lack the qualities of a killer app—one that’s built on real business cases and can energize managers. Sector- and company-specific business factors are powerful enablers (or enemies) of successful data efforts. That’s why it’s crucial to give planning a second dimension, which seeks to balance the need for affordability and speed with business realities (including easy-to-miss risks and organizational sensitivities).
To understand the costs of omitting this step, consider the experience of one bank trying to improve the performance of its small-business underwriting. Hoping to move quickly, the analytics group built a model on the fly, without a planning process involving the key stakeholders who fully understood the business forces at play. This model tested well on paper but didn’t work well in practice, and the company ran up losses using it. The leadership decided to start over, enlisting business-unit heads to help with the second effort. A revamped model, built on a more complete data set and with an architecture reflecting differences among various customer segments, had better predictive abilities and ultimately reduced the losses. The lesson: big-data planning is at least as much a management challenge as a technical one, and there’s no shortcut in the hard work of getting business players and data scientists together to figure things out.
At a shipping company, the critical question was how to balance potential gains from new data and analytic models against business risks. Senior managers were comfortable with existing operations-oriented models, but there was pushback when data strategists proposed a range of new models related to customer behavior, pricing, and scheduling. A particular concern was whether costly new data approaches would interrupt well-oiled scheduling operations. Data managers met these concerns by pursuing a prototype (which used a smaller data set and rudimentary spreadsheet analysis) in one region. Sometimes, “walk before you can run” tactics like these are necessary to achieve the right balance, and they can be an explicit part of the plan.
At a health insurer, a key challenge was assuaging concerns among internal stakeholders. A black-box model designed to identify chronic-disease patients with an above-average risk of hospitalization was highly accurate when tested on historical data. However, the company’s clinical directors questioned the ability of an opaque analytic model to select which patients should receive costly preventative-treatment regimes. In the end, the insurer opted for a simpler, more transparent data and analytic approach that improved on current practices but sacrificed some accuracy, with the likely result that a wider array of patients could qualify for treatment. Airing such tensions and trade-offs early in data planning can save time and avoid costly dead ends.
Finally, some planning efforts require balancing the desire to keep costs down (through uniformity) with the need for a mix of data and modeling approaches that reflect business realities. Consider retailing, where players have unique customer bases, ways of setting prices to optimize sales and margins, and daily sales patterns and inventory requirements. One retailer, for instance, has quickly and inexpensively put in place a standard next-product-to-buy model3 for its Web site. But to develop a more sophisticated model to predict regional and seasonal buying patterns and optimize supply-chain operations, the retailer has had to gather unstructured consumer data from social media, to choose among internal-operations data, and to customize prediction algorithms by product and store concept. A balanced big-data plan embraces the need for such mixed approaches.

3. Ensuring a focus on frontline engagement and capabilities

Even after making a considerable investment in a new pricing tool, one airline found that the productivity of its revenue-management analysts was still below expectations. The problem? The tool was too complex to be useful. A different problem arose at a health insurer: doctors rejected a Web application designed to nudge them toward more cost-effective treatments. The doctors said they would use it only if it offered, for certain illnesses, treatment options they considered important for maintaining the trust of patients.
Problems like these arise when companies neglect a third element of big-data planning: engaging the organization. As we said when describing the basic elements of a big-data plan, the process starts with the creation of analytic models that frontline managers can understand. The models should be linked to easy-to-use decision-support tools—call them killer tools—and to processes that let managers apply their own experience and judgment to the outputs of models. While a few analytic approaches (such as basic sales forecasting) are automatic and require limited frontline engagement, the lion’s share will fail without strong managerial support.
The aforementioned airline redesigned the software interface of its pricing tool to include only 10 to 15 rule-driven archetypes covering the competitive and capacity-utilization situations on major routes. Similarly, at a retailer, a red flag alerts merchandise buyers when a competitor’s Internet site prices goods below the retailer’s levels and allows the buyers to decide on a response. At another retailer, managers now have tablet displays predicting the number of store clerks needed each hour of the day given historical sales data, the weather outlook, and planned special promotions.
But planning for the creation of such worker-friendly tools is just the beginning. It’s also important to focus on the new organizational skills needed for effective implementation. Far too many companies believe that 95 percent of their data and analytics investments should be in data and modeling. But unless they develop the skills and training of frontline managers, many of whom don’t have strong analytics backgrounds, those investments won’t deliver. A good rule of thumb for planning purposes is a 50–50 ratio of data and modeling to training.
Part of that investment may go toward installing “bimodal” managers who both understand the business well and have a sufficient knowledge of how to use data and tools to make better, more analytics-infused decisions. Where this skill set exists, managers will of course want to draw on it. Companies may also have to create incentives that pull key business players with analytic strengths into data-leadership roles and then encourage the cross-pollination of ideas among departments. One parcel-freight company found pockets of analytical talent trapped in siloed units and united these employees in a centralized hub that contracts out its services across the organization.
When a plan is in place, execution becomes easier: integrating data, initiating pilot projects, and creating new tools and training efforts occur in the context of a clear vision for driving business value—a vision that’s unlikely to run into funding problems or organizational opposition. Over time, of course, the initial plan will get adjusted. Indeed, one key benefit of big data and analytics is that you can learn things about your business that you simply could not see before.
Here, too, there may be a parallel with strategic planning, which over time has morphed in many organizations from a formal, annual, “by the book” process into a more dynamic one that takes place continually and involves a broader set of constituents.4 Data and analytics plans are also too important to be left on a shelf. But that’s tomorrow’s problem; right now, such plans aren’t even being created. The sooner executives change that, the more likely they are to make data a real source of competitive advantage for their organizations.

About the authors

Stefan Biesdorf is a principal in McKinsey’s Munich office, David Court is a director in the Dallas office, and Paul Willmott is a director in the London office.
The authors would like to acknowledge the contributions of Toos Daruvala, Amit Garg, and David Kang to the development of this article.