Showing posts with label My Nerdie Bookshelf. Show all posts
Showing posts with label My Nerdie Bookshelf. Show all posts

Wednesday, August 05, 2015

My Nerdie Bookshelf - "Linked Data - Structured data on the web" by David Wood, Marsha Zaidman, Luke Ruth and Michael Hausenblas

This book has been a bit of a disappointment to me, the first one I had from Manning Publications.

Despite being published in 2014 you have the impression that the information provided here are stale. Only in the final chapter ("The evolving Web") a comprehensive, well written and updated viewpoint on Linked Data (and the Semantic Web) is provided, although in concise form.

The foreword by Tim Berners-Lee and the collaboration with Michael Hausenblas lured me to blind purchase the book. In particular I was looking for insights in what, in my viewpoint, is a powerful use case for Linked Data which hasn't been addressed enough, that is Semantic Enterprise Data Integration, hoping to get, as it is common for Manning books, a lot of advanced technical information. In particular I was, and still am, looking for technical advice, integration patterns and product reviews that can guide me in using Semantics to effectively interconnect enterprise data silos.

The book on the other hand revolves around a different perspective, those of a data publisher, with little (if any) notion of the technology behind Linked Data. It presents therefore all the basic concepts at a quite simple level.
This is of course a legit editorial choice but what annoyed me the most was the fact that the information provided are often outdated. No mention on JSON-LD or to the Linked Data Platform principles; CKAN, a widely used platform for creating open data repositories, is just cited but only in connection to the DataHub site. Moreover, the motivations, advantages, pros and cons of working on Linked Data are presented in a very basic, if not superficial, way.

The mention of Callimachus, the "Linked Data application server" created by the authors, left me unimpressed as well, even if it is correct to say that it has been used in interesting projects.

I must admit that I am biased and might sound arrogant (sorry if this is the case): at the end of the day I've been working on these topics for 5 years. The fact is that this book could have been appealing to beginners if only could present more up-to-date information and more detailed use cases. Linked Data looks like it was written in 2010: it could make sense to publish it in 2011, not, as it was the case, in 2014.

Sunday, September 14, 2014

My Nerdie Bookshelf - "Big Data" by Nathan Marz and James Warren

Undoubtedly Big Data is the hype of 2014: if on the one hand the availability of cheap cloud resources, together with the huge increase of data sources, have paved the way for the incredible rise of interest in this subject, it is also true that behind it there is also a lot of marketing.

Real use cases are abundant but Big Data has become a buzzword used now for everything regarding data analytics. Interesting viewpoints on that can be read on selected articles like "'Big Data' Is One Of The Biggest Buzzwords In Tech That No One Has Figured Out Yet" and "If Big Data Is Anything at All, This Is It". Moreover often the term "Big" is improperly used. Definitely 100 Gb of overall data, that can be processed in RAM on affordable machines (see for example this offer from Hetzner), can't be defined Big Data...

The strongest point of the Big Data book I'm reviewing here has been for me its ability to present a clear definition of the problem. Since the hype, I wanted to study more about Big Data and chose this Nathan Marz book thanks to the very good reputation of its author (well, co-author to say the truth): Marz was a technical architect at Twitter and founded several exciting open source projects (one for all Storm, a sort of ultra scalable, high-performance service bus).

Marz presents here his vision and recipes to deal with the Big Data problem, namely the Lambda Architecture. It is definitely a high end solution, to be used when great bunch of data must be processed with the lowest latency as possible. For this purpose the Lambda Architecture consists of two layers, the batch layer and the speed layer, the latter, as the name implies, to process the most recent data.

Albeit the book is very practical and describes directly the technical solutions and the open source technologies the Lambda Architecture is based on, it also expresses clearly the characteristic of Big Data processing. At the heart of it all there is the idea that analytics is just a "function" of all the available data and that the master dataset is immutable: new information should be only accumulated and never deleted. This way it is possible to reprocess the whole universe of information when needed. This allows several advantages like the ability to easily correct errors introduced in previous processing and the ability to elaborate new indicators or to refine existing ones, on the whole set of information. These data in fact must be processed to produce (or to enrich if processing is incremental) indexes used for queries (eg for business intelligence): in the author’s words, “to make queries on precomputed views rather than directly on the master dataset”.

The key technology here is of course Hadoop, both for the distributed management of big data and for processing through map-reduce powered algorithms. Quite interesting is the chapter about modeling, that presents a data schema that revolves around atomic facts, logically linked to form a graph of structures: this at the end is not too dissimilar to the star/snowflakes schemas that can be found in traditional OLAP data warehouses.

For the real time layer Marz proposes the use of the aforementioned Storm, in this case to update short-lived indexes, destined to be removed after a few hours when the batch layer is able to ingest the recent data they were based on.

While there is a hint of theory here and there, Big Data is very much practical and highly focused in presenting the Lambda Architecture (a much more suitable title should have been A reference architecture for Big Data processing). I wouldn't advice the book to developers totally new to the Big Data ecosystem: it presents in fact several specific technical solutions (Thrift, Pail, JCascalog, Cassandra, ElephantDB, Trident,…) that should help and simplify the problems described, but can be probably more appreciated only to those already skilled on the subject. In my viewpoint in fact it would be beneficial first to get the hands dirty on all the main technologies behind (Hadoop to start with and Storm just in case) before directly jumpstart to embrace the final Lambda Architecture as it is presented here.

Overall anyway I found the book interesting and stimulating: the fact that MEAP books can be easily bought with big discounts (just follow the Manning Twitter account to get promotional coupons) is also a plus worth mentioning here.

More information on Big Data can be found on the Manning web site.


Monday, July 07, 2014

My Nerdie Bookshelf - "Getting things done - The art of stress-free productivity" by David Allen

I don't remember exactly how I heard about this book but it immediately caught my attention when I heard that it suggests productivity tricks based on the use of to-do lists. I was and still am a great fan of lists: they really helped me to untangle the huge mass of tasks and problems I have dealt with in the last ten years. I'm using lists for almost everything, including when I'm preparing the luggage for a trip.

I must admit that the author's method is quite interesting and effective. To say that it is simply based on lists is a huge understatement. David Allen proposes a process to deal with everyday's activities in which for every task that must be faced a "next action" is defined. Lists and calendars serve for keeping the mind free, instead of trying to maintain schedules, appointments, deadlines, notes and things to do all in one's memory.

Most of the advices proposed in Getting things done are purely based on common sense, but they are strangely often forgotten or dismissed, especially in work environments.
I've learnt some interesting few tricks from this book. Three are worth mentioning here, that is: the 2 minutes rule (if you need less than 2 minutes to do something don't defer it and do it now), the importance of always identifying a next action for each task and defining in meetings a clear purpose at the beginning and a list of actions (and responsibilities) at the end.

The book has been initially published in 2001 and there isn't much emphasis on the exploitation of technology (besides some random mentions of PDAs - that was way before smartphones! - and personal productivity features of desktop softwares like Outlook or Lotus Notes). This should not be considered detrimental of the book: its value is in the process, not in the suggested (low tech) implementation.

Despite these good things there are also some annoying facts about Getting things done. To start with, its messianic style, the constantly repeated mantra that your life can deeply change if you follow the advices of this book. Moreover while the method can really apply to everyone (and, to say the truth, this is often repeated throughout the pages) the examples in Getting things done refer always to a vip audience. All this gives the impression that David Allen is constantly promoting his activity as a consultant: this is legit of course, but it is sincerely annoying to get this feeling in a book you've bought.

Finally, despite the book is not very long, often concepts are repeated over and over again. You could actually get the same value from it even if it had 40% less pages.

Despite these flaws I deeply suggest the reading of Getting things done to those who feel in trouble to keep the pace of business and personal events. I honestly learnt some useful strategies here, which at the end is the only thing that matters in a book like this.

Getting things done - The art of stress-free productivity, David Allen, Penguin Books 2001

Source Wikipedia

Introducing My Nerdie Bookshelf

I love reading: this is a thing in the family, since between my wife, our 10 years old son and me our house is really flooded with books.

I enjoyed keeping track of the books I read, both in paper and electronic form, but I never liked to mix my leisure time with my professional reading. Therefore even at home I try to keep separate my technical bookshelf (My nerdy bookshelf) with the other (huge…) ones used to store everything, from gothic novels to italian thriller, from art books to comics.

I enjoy using Anobii to keep track of the “leisure time” books I read and buy, but, for what I’ve said before, I don’t feel like using it for my nerdy reading as well. At the same time, I do really feel the need to write some comments after I’ve read a professional book. That’s why I fancied about reactivating this very neglected blog to start the My Nerdy Bookshelf series of posts. Here I’ll report comments for the technical books I read, mostly related to my present IT profession of Project Manager and Software Architect (Ok, let’s also add Entrepreneur, since I’m an associate of Net7 srl).

I’m going to write these posts mostly for myself and of course if they can be valuable for others I’ll be more than happy. At the same time it’s going to be a personal thing, not a “professional” blogger activity. I’ll report my personal impressions, with humility and honesty: my 2 cents, hopefully with attitude!