Big Data
Using Themes for Enhanced Problem Solving
Thematic Analysis is a powerful qualitative approach used by many consultants. It involves identifying patterns and themes to better understand how and why something happened, providing context for other quantitative analyses. It can also be used when developing strategies and tactics because of its “cause and effect” nature.
Typical analysis tends to be event-based. Something happened that was unexpected. Some type of triggering or compelling event is sought to either stop something from happening or to make something happen. With enough of the right data, you may be able to identify patterns that help predict what will happen next based on past events. This data-based understanding may be simplistic or incomplete, but often it is sufficient.

But people are creatures of habit. If you can identify and understand those habits and place them within the context of a specific environment that includes interactions with others, you may be able to identify patterns within the patterns. Those themes can be much better indicators of what may or may not happen than the data itself. They become better predictors of things to come and can help identify more effective strategies and tactics to achieve your goals.
This approach requires that a person view an event (desired or historical) from various perspectives to help understand:
- Things that are accidental but predictable because of human nature.
- Things that are predictable based on other events and interactions.
- Things that are the logical consequence of a series of events and outcomes.
Aside from the practical implications of this approach, I find it fascinating relative to AI and Predictive Analysis.
For example, you can proactively monitor data, activities, and patterns by understanding recurring themes and triggers. That provides actionable intelligence that can be automated and incorporated into a larger system. Machine Learning and Deep Learning can analyze tremendous volumes of data from various sources in real-time.
Combine that with Semantic Analysis (and products like AllegroGraph) to harness the power of language through taxonomies and ontologies. The monitoring system would better understand what is happening (not just be aware of it), possibly determine how and why it happened, and result in even more focused and accurate predictions. Finally, include spatial and temporal data such as IoT, metadata from photographs, etc., and you should be able to view something as though you were very high up – providing the ability to “see” what is on the path ahead. It is obviously not that simple, but it is exciting.
This approach provides a multi-dimensional view of events and their causality, giving you the means to identify and prevent complex problems before they negatively affect a business.
Keeping these thoughts in mind will help you see details others have missed. Better tools can make for better analysis, better strategies, and better outcomes. Who wouldn’t want that?
The Coming Changes to Manufacturing
Recently, I spoke with someone on a team analyzing ways to “mitigate the risk of exclusive manufacturing in China” without fully divesting their business interests in a growing and potentially lucrative market. This bifurcation exercise got me thinking about how many other companies are evaluating their supply chain relationships, inventory management, and the predictability of their cost of goods sold.

In the mid-1990s, I had done a lot of work with the MK manufacturing software that ran on the Ingres database. Some issues were performance-related and fixed by database tuning; some were fixed by using average costs instead of a full Bill of Materials (BOM) explosion with dozens of screws in a window; but some were more interesting and more business-focused.
After NAFTA became law, one manufacturer built a facility in Mexico and started manufacturing a few basic but important parts. When I arrived as a Consultant, the main problem they faced was a reject rate of roughly 20% and additional related QA costs. My suggestion was to treat this part (a single piece of steel, like the rotor from a disk brake system) as a component and build in the cost of both scrap and QA. They could then benchmark the costs against other suppliers in an apples-to-apples comparison to determine if they really saved money. That approach worked well for them.
While that approach helped manage costs, it did not address the timeliness of orders or lead time required – important aspects of Just-in-Time (JIT) manufacturing. Additionally, it should be possible to estimate shipping costs by considering changes in petroleum costs or anticipated changes in demand or capacity.
Systems out there claim to estimate the cost and availability of commodities based on various global factors and leading indicators. It is tricky, to say the least, and we can’t anticipate an event like a pandemic. But companies that manage their inventory and production risk best will likely be the ones that succeed in the long run. They will become the most reliable suppliers and have increased profits to invest in further growth and improvement.
The next 2-3 years will be very interesting due to technological advances (especially AI) and geopolitical changes. Those companies that embrace change and focus on real transformation will likely emerge as the new leaders in their segments by 2025.
Blockchain, Data Governance, and Smart Contracts in a Post-COVID-19 World
The last few months have been very disruptive to nearly everyone across the globe. There are business challenges galore, such as managing large remote workforces – many of whom are new to working remotely and managing risk while attempting to conduct “business as usual.” Unfortunately, most businesses’ systems, processes, and internal controls were not designed for this “new normal.”
While there have been many predictions around Blockchain for the past few years, it is still not widely adopted. We are beginning to see an uptick in adopting Supply Chain Management Systems for reasons that include traceability of items – especially food and drugs. However, large-scale adoption has been elusive to date.

I believe we will soon see major shifts in mindset, investment, and effort toward modern digital technology driven by Data Governance and Risk Management. I also believe that this will lead to these technologies becoming easier to use via new platforms and integration tools, which will lead to faster adoption by SMBs and other non-enterprise organizations, and that will lead to the greater need for DevOps, Monitoring, and Automation solutions as a way to maintain control of a more agile environment.
Here are a few predictions:
- New wearable technology supporting Medical IoT will be developed to help provide an early warning system for disease and future pandemics. That will fuel innovations across industries, including Biotech and Pharma.
- Blockchain can provide data privacy, ownership, and provenance to ensure the data’s veracity.
- New legislation will be created to protect medical providers and other users of that data from being held liable for missing information or trends that could have saved lives or avoided other negative outcomes.
- In the meantime, Hospitals, Insurance Providers, and others will do everything possible to mitigate the risk of using Medical IoT data, which could include Smart Contracts to ensure compliance (assuming a benefit is provided to the data providers).
- Platforms may be created to offer individuals control over their own data, how it is used and by whom, ownership of that data, and payment for the use of that data. I wrote about this in 2013.
- Data Governance will be taken more seriously by every business. Today, companies talk about Data Privacy, Data Security, or Data Consistency, but few have a strategic end-to-end systematic approach to managing and protecting their data and their company.
- Comprehensive Data Governance will become a driving and gating force as organizations modernize and grow. Even before the pandemic, there were growing needs due to new data privacy laws and concerns around areas such as the data used for Machine Learning.
- In a business environment where more systems are distributed, the risk of data breaches and Cybercrime Increases. That must be addressed as a foundational component of any new system or platform.
- One or two Data Integration Companies will emerge as undisputed industry leaders because of their capabilities in MDM, Data Provenance and Traceability, and Data Access (an area typically managed by application systems).
- New standardized APIs akin to HL7 FHIR will be created to support a variety of industries as well as interoperability between systems and industries. Frictionless integration of key systems becomes even more important than it is today.
- Anything that can be maintained and managed in a secure and flexible distributed digital environment will be implemented to allow companies to quickly pivot and adapt to new challenges and opportunities on a global scale.
- Smart Contracts and Digital Currency Payment Processing Systems will likely be core components of those systems.
- This will also foster the growth of next-generation Business Ecosystems and more dynamic collaborations.
- Ongoing compliance monitoring, internal and external, will likely become a priority (“trust but verify”).
All in all, this is exciting from a business and technology perspective. Most companies must review and adjust their strategies and tactics to embrace these concepts and adapt to the coming New Normal.
The steps we take today will shape what we see and do in the coming decade, so it is important to get this right quickly, knowing that whatever is implemented today will live, evolve, and hopefully improve over time. Don’t wait for perfection, as the risks are too high.
Good Article on Why AI Projects Fail

Today I came across this very good article focused on lessons learned, which could help anyone interested in these topics. It included a good mix of non-technical problems.
This is the link to the article, along with my commentary on the Top 3 items listed: https://www.cio.com/article/3429177/6-reasons-why-ai-projects-fail.html
Item #1:
The article discusses how the “problem” being evaluated was misstated using technical terms. At least some of these efforts are conducted “in a vacuum.” Given the cost and strategic importance of getting these early-adopter AI projects right, that surprised me.
In Sales and Marketing, you start the question, “What problem are we trying to solve?” and evolve that to, “How would customers or prospects describe this problem in their own words?” Without that understanding, you can neither vet the solution initially nor quickly qualify the need for it when speaking with customers or prospects. That leaves room for error when transitioning from strategy to execution.
More collaboration with Business likely would have helped. This was touched on at the end of the article under “Cultural challenges,” but the importance seemed to be downplayed. Lessons learned are valuable – especially when you are able to learn from the mistakes of others. This should have been called out early as a major lesson learned.
Item #2:
This second area had to do with the perspective of the data, whether that was the angle of the subject in photographs (overhead from a drone vs horizontal from the shoreline) or the type of customer data evaluated (such as from a single source) used to train the ML algorithm.
That was interesting because assumptions may have played a role in overlooking other aspects of the problem, or the teams may have been overly confident they could get the right results with the data available. In the examples cited, those teams identified the problems and took corrective action. A follow-up article describing the process used to determine the root cause in each case would be very interesting.
As an aside, from my perspective, this is why Explainable AI is so important. Sometimes, you just don’t know what you don’t know (the unknown unknowns). Understanding why and on what the AI is basing its decisions should help provide better-quality curated data up front, as well as identify potential drifts in the wrong direction while it is still early enough to make corrections without impacting deadlines or deliverables.
Item #3:
This didn’t surprise me, but it should be a cause for concern as advances are made at faster rates and organizations race to be first to market with an AI-based competitive advantage, potentially with less validation than ideal. The last paragraph under ‘Training data bias’ stated that based on a PWC survey, “only 25 percent of respondents said they would prioritize the ethical implications of an AI solution before implementing it.“
Bonus Item:
The discussion about the value of unstructured data was very interesting, especially when you consider:
- The potential for NLU (natural language understanding) products in conjunction with ML and AI.
- This is a great NLU-pipeline diagram from North Side Inc. in Canada, one of the pioneers in this space.
- The importance of semantic data analysis relative to any ML effort.
- The incredible value that products like MarkLogic’s database or Franz’s AllegroGraph provide over standard Analytics Database products.
- I personally believe the biggest exception to this assertion will be GPU databases (like OmniSci) that easily handle streaming data, can accomplish extreme computational feats well beyond traditional CPU-based products, and have geospatial capabilities that add an extra dimension of insight to the problem being solved.
Update: This is a link to a related article that discusses trends in areas of implementation, important considerations, and the potential ROI of AI projects: https://www.fastcompany.com/90387050/reduce-the-hype-and-find-a-plan-how-to-adopt-an-ai-strategy
This is an exciting space that will grow significantly over the next 3-5 years. The more information, experiences, and lessons learned are shared, the better it will be for everyone.
The Unsung Hero of Big Data
Earlier this week, I read a blog post regarding the recent Gartner Hype Cycle for Advanced Analytics and Data Science, 2015. The Gartner chart reminded me of the epigram, “Plus ça change, plus c’est la même chose” (asserting that history repeats itself by stating the more things change, the more they stay the same).
To some extent, that is true, as you could consider today’s Big Data as a derivative of yesterday’s VLDBs (very large databases) and Data Warehouses. One of the biggest changes, IMO, is the shift away from Star Schemas and practices implemented for performance reasons, such as aggregation of data sets, using derived and encoded values, using surrogate and foreign keys to establish linkage, etc. Going forward, it may not be possible to have that much rigidity and still be as responsive as needed from a competitive perspective.
There are many dimensions to big data: A huge sample of data (volume), which becomes your universal set and supports deep analysis as well as temporal and spatial analysis; A variety of data (structured and unstructured) that often does not lend itself to SQL based analytics; and often data streaming in (velocity) from multiple sources – an area that will become even more important in the era of the Internet of Things. These are the “Three V’s” people have talked about for the past five years.
Like many people, my interest in Object Database technology initially waned in the late 1990s. That is, until about four years ago, when a project at work led me back in this direction. As I dug into the various products, I learned they were alive and doing well in several niche areas. That finding led to a better understanding of the real value of object databases.
Some products try to be “All Vs to all people,” but generally, what works best is a complementary and integrated set of tools working together as a service within a single platform. It makes a lot of sense. So, back to object databases.
One of the things I like most about my job is the business development aspect. One of the product families I’m responsible for is Versant. With the Versant Object Database (VOD – high performance, high throughput, high concurrency) and Fast Objects (great for embedded applications like kiosks). I’ve met and worked with brilliant people who have created amazing products based on this technology. Creative people like these are fun to work with, and helping them grow their business is mutually beneficial. Everyone wins.
An area where VOD excels is with the near real-time processing of streaming data. The reason it is so adept at this task is the way that objects are mapped out in the database. They do so in a way that essentially mirrors reality. So, optionality is not a problem – no disjoint queries or missed data, no complex query gyrations to get the correct data set, etc. Things like sparse indexing are not a problem with VOD. This means that pattern matching is quick and easy, as well as more traditional rule and look-up validation. Polymorphism allows objects, functions, and even data to have multiple forms – something else that mirrors real life (just think about the variations of a peripheral device called a “printer”).
VOD and products like it do more by allowing data to be more, which is ideal for environments where change is the norm, such as: Cyber Security; Fraud Detection; Threat Detection; Logistics; and Heuristic Load Optimization. In each case, performance, accuracy, and adaptability are the key to ongoing success.
The ubiquity of devices generating data today, combined with the desire for people and companies to leverage that data for commercial and non-commercial benefit, is very different than what we saw 10+ years ago. Products like VOD are working their way up that “Slope of Enlightenment” because there is a need to connect the dots better and faster – especially as the volume and variety of those dots increases.
It is not a “one size fits all” solution, but it is often the perfect tool for complex data. More importantly, it is another tool to use in an ever-expanding data ecosystem.
These are indeed exciting times!
