Big Data Solution Architect Resume
Hopkins, MN
SUMMARY:
- A 15+ years experienced Data Integration and Business Intelligence Professional, worked more than a decade on various successful data integration and Business Intelligence project implementations to create better value for client's data. Worked on industry leading data integration, visualization tools and technologies in bringing data from different sources systems together. Produced interactive and insightful visualizations to assist business users to make better decisions every day. Presently working with structured, unstructured, relational and various other formats data to build strong data pipelines with Hadoop Eco System, HDFS, Hive, Sqoop, Spark and Kafka. Helped Businesses in setting up their Hortonworks Data Cloud - HDP Services on AWS EC2 instances successfully. Worked on Hortonworks Data Cloud - Controller Service to setup HDP clusters and enabled users to model and analyze large sets of Data.
- Led and executed client’s Business Intelligence initiatives, used propelling innovative ideas and techniques to save cost and increase user adaptability in the BI space.
- Developed streamlined, efficient, and robust data architectures to meet business needs for data migration, information delivery, reporting and analytics.
- Worked on Data-as-a-Service and Platform-as-a-service offerings to make easy and cost-effective processing of large-scale data sets in the cloud.
- Worked on Hortonworks and Cloudera Hadoop Distributions to develop massive and robust data pipelines. Worked on Data Ingestion, Data Transformation and Data Attribution to build Hybrid Data Lake to answer valuable business questions.
- Having a great deal of expertise in dealing with relational, No-Sql and Graph databases.
- Worked on Continuous Integration, Deployment and DevOps tools and technologies to help and enforce automation, to ensure smooth transition between environments and so to improve overall quality of any project running on the big data platform.
- Hadoop and Cloud Eco Systems integration with DevOps tools like Jenkins, TeamCity, Ansible, Git, Maven, Chef, Puppet, Docker and Kubernetes to monitor each big data component, projects and quality (delivery, performance & SLA, logs).
- Passionate about storytelling, data Integration, Data Mining, Business Intelligence, Bigdata, Hadoop Distributions on cloud, Data Science and machine learning.
- Skilled in providing unified solutions with central focus on business value, customer satisfaction, and future vision.
- Strong programming skills in Python and Scala to build efficient and robust data pipelines.
- Project Management experience and Subject Matter knowledge to solve common and complex business issues within established guidelines.
- Proposed and implemented BI Best Practices within a number of projects that has ultimately saved on cost.
TECHNICAL SKILLS:
Hadoop and Big Data: HDFS, Hive, Map Reduce, Sqoop, Spark, Kafka, Flume, Flink, YARN, Oozie, Zookeeper, Ambari, Cloudera and Hortonworks Distributions
Data Integration and ETL Tools: Informatica, IBM Data Stage, SSIS, Talend and SnapLogic
Business Intelligence Tools: SAP Business Objects 4.1/3.X, Tableau, SSRS, SSAS, Qlikview, Cognos, and TIBCO Spotfire
Programming Languages: Python, Scala, Groovy, Shell and Batch programming
Databases: Netezza, DB2, Oracle, SQL Server, Sybase, Teradata, SAP HANA, Hbase, MongoDB, Cassandra, Neo4J
Data Modelling Tools: Embarcadero, ERwin and Power Designer
Cloud and Data Science Technologies: AWS Ec2, S3, RedShift, Elastic Search. EMR, Kinesis, RDS, Dynamo DB, Lambda and Quicksight, Microsoft Azure, Google Cloud Platform, Data Lake, SAS and R Programming.
DevOps Tools: Jenkins, TeamCity, SonarQube, Ansible, Git, Maven, Terraform, Chef, Puppet, Docker and Kubernetes
Versioning Tools: Git, Github, Bit Bucket and Subversion
Methodologies: Agile and Lean Methodologies, SDLC, Jira, IBM RTC
Industries: Health Insurance, Banking, Retail, Manufacturing, Finance and Insurance
Operating Systems: Linux, Windows and UNIX
PROFESSIONAL EXPERIENCE:
Confidential, Hopkins, MN
Big Data Solution Architect
- Meeting business users to understand their data and application requirements. Provide solutions, recommend technical stack of tools and technologies for development and precisely address business use cases.
- Consult with senior leadership and business leaders to provide demos on developed work products on regular intervals. Receive feedback from business and work with development team to add more business value to deliverables.
- Analysing multiple source systems like Relational, Real-time, Structured and Unstructured sources and work on development of batch and real-time data pipelines to get data into Datalake.
- Developing Data pipelines for ingestion of various formats of data into Hive/Impala/kudu tables using HDFS, Sqoop, Spark, Scala, Python, MFT, Oozie and Kafka.
- Extraction, loading and Transformation of business critical data from SAP source systems like SAP IBP, HANA, ECC and HR.
- Developing data pipelines using Stream Sets, Apache NI-FI and AWS Glue.
- Developing data pipelines to process data coming from AZURE enterprise APIs by making CURL calls and processed data with near real time.
- Developing data streaming applications to process sensor information on near real time basis, used machine learning and data science techniques to predict critical metrics and key figures.
- Working on Kafka configuration of producers and listeners for processing of streaming data. Developed data pipelines to process and load streaming data into Hive internal tables and Cassandra.
- Development of Oozie/rundeck workflows to run and monitor all ingestion processes and schedules on production.
- Implemented role based security using control tables for users, user groups, process IDs and AD groups using Impala tables.
Confidential, Irving, TX
Big Data Architect
- Interact with business users and senior leadership to gather requirements and analysis of multiple source systems data like Relational, Real-time, Structured and Unstructured Data.
- Build Data pipelines for ingestion of various formats (Ex: Json, csv, txt ) of data into Hive internal tables using HDFS, Sqoop, Spark, Scala, MFT, Oozie and Kafka.
- Extraction, Loading and Transformation of business critical data (Ex: Collect and Process Stream logs, csv to json, data wrangling) using Apache Ni-Fi, Spark SQL, Spark Streaming.
- Evaluated and working on Amazon Glue as an ETL tool to process business critical data into aggregated tables in Hive Cloud. Deployed and Development in Bigdata applications like Spark, Hive, Kafka and flink in AWS cloud.
- Developed multiple APIs using Spark MLib, Spark SQL and Spark Streaming, which are capable to process millions of records each day.
- Used Data Science algorithms (Ex: RegEx, Support Vector Machines, Logistic Regression and Naïve Bayes) for development of prediction models in Spark Scala.
- Worked on prediction metrics like precision, sensitivity, accuracy, fmeasure, tp, tn, fp, fn and predicted counts and developed in sparkML and Scala.
- Worked on Kafka configuration of producers and listeners for processing of streaming data. Developed data pipelines to process and load streaming data into Hive internal tables and Cassandra.
- Developed Oozie workflows to run and monitor all ingestion processes and schedules.
- Heavily worked on creating training data to feed into Machine Learning models for accurate prediction of data.
- Worked on Hive internal and external tables for data modelling, processing and storage.
- Proficient in UNIX Scripting for Data Processing and Automation. Evaluated H2O.ai platform for the project and started training the models with supervised and unsupervised learning for accurate predictions.
- Worked on Tableau to create Dashboards and Scorecards to visualize business critical metrics.
- Performance tuning of spark applications by configuring the driver memory, executor memory, increasing the cores and queues for spark jobs with limitation.
- Building the process groups and processes in the Nifi to pull the files from the various servers and placing the files in the HDFS and components to convert it into JSON and evaluate and store the file information in the file tracker and in Kafka topics.
- Worked on AWS Cloud to convert all on premise, existing processes and data bases to AWS Cloud.
- Successfully implemented DevOps Methodologies and Technologies (Python, Jenkins, Nexus and Docker) in the project for assuring Continuous Integration and Continuous Delivery on AWS cloud.
- Worked on CI/CD solution, using Git, BitBucket, Jenkins, Docker, Nexus and Kubernetes to setup and configure Big data architecture on AWS cloud platform.
Confidential, PLANO, TX
Sr. Big Data Engineer / Analyst
- Gather requirements and analysis of multiple source systems data like Relational, Click Strems and portals
- Build Data pipelines for Extraction, Loading and Transformation of business critical data.
- Migrated from IBM Big Query and GPFS to Horton works Eco system, Utilized HDFS, Python scripting, Hive, Sqoop, Spark and Kafka for building new age data pipelines for bulk ingestion of relational and non-relational data.
- Successfully moved Data Lake and ETL processes from GPFS to HDFS without any issues.
- Handle large volumes of data and loaded into Hive tables and Enterprise DWH on daily, weekly and monthly basis. Automated the batch load and real time loading processes with Shell Scripting and Control M.
- Performance tuning of spark applications by configuring the driver memory, executor memory, increasing the cores and queues for spark jobs with limitation.
- Building the process groups and processes in the Nifi to pull the files from the various servers and placing the files in the HDFS and components to convert it into JSON, evaluated and stored the file information in the file tracker and in Kafka topics.
- Migrated IBM DataStage 11.5 jobs to Spark ETL jobs, aggregated the daily data and load into EDW on weekly and monthly basis.
- Develop Automated Quality checks with UNIX shell scripts and reusable Source prep and Loading jobs to verify and reconcile data during data loads.
- Developed Sqoop wrappers with python scripting to FTP the heavy volumes of data between Data Lake and Data Warehouse.
- Worked on data visualization tools and technologies (SAP Business Objects and Tableau) to showcase aggregated data in Dashboards and interactive visualizations.
- Participate in project road map discussions and handle daily Scrum meetings in 4 week iterations.
- Successfully implemented DevOps Methodologies and Technologies (Python, Jenkins, Nexus and Docker) in the project for assuring Continuous Integration and Continuous Delivery on AWS cloud.
- Worked on R programming, SCALA and AWS S3, Redshift, Ec2 and used them in project for more effective data mining and data discovery.
- Design and implement CI/CD solution, using Git, BitBucket, Jenkins, Docker, Nexus and Kubernetes.
Confidential
Big Data Analyst
- Responsible for translating architectural specifications from Solution and/or Enterprise Architecture and created a technical design document.
- Created Data Lake by extracting data from real time customer actions into HDFS. Implement data ETL with Map-Reduce in Spark SQL.
- Build Data pipeline upon AWS EC2 with S3, EMR, Elastic Search, Kinesis, Dynamo DB and Redshift. Manage distributed cluster as AWS admin role.
- Compose distribution system based on Apache Spark framework. Implement Spark Streaming application with Scala for customer request streaming management and modeling.
- Built Machine Learning Pipeline and ETL processing based on Spark distributed platform. Applied Spark MLlib and Random Forest algorithm with Scala to detect anomaly data.
- Ingestion multiple data source with Kinesis Stream, and loading data into S3 bucket. Install, deploy, update and maintenance Cloudera Big Data ecosystem in EC2 instance and Designed Redshift database for data science and BI team.
- Developed a string similarity spark application for matching company names from multiple source using edit distance algorithms and Spark Scala API; tuned the spark jobs for optimal efficiency by increasing parallelism and reducing shuffles; implemented a spelling check engine using Python NLP to automatically correct misspelled company names.
- Created and processed RDD's and Data Frames using SparkSQL. Design and develop Shell Scripts, Pig Scripts and Hive Scripts and MapReduce jobs.
- Hive queries and partitions to store the data in internal table. Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Performance tuning of Spark Applications, analyzing various dependencies, storage levels, resource tuning and memory management.
- Converted large sets of IBM DataStage jobs to Informatica power center mappings to facilitate IBM DataStage and Sybase retirement.
- Architected, estimated and developed the complete ETL migration project by creating Informatica mappings and workflows.
- Modelled, designed and developed new Enterprise data warehouse on Teradata platform and retired the old warehouse on Sybase.
- Converted large SAP Business Objects universes and complex reports from SAP BO 3.1 to latest 4.1 version, as well as .UNV universes on Sybase to .UNX universes on Teradata.
- Designed, developed and migrated interactive dashboards and Analytics with Tableau Desktop Server using calculations and parameters
- Worked on a PoC using SAS and R programming to implement deep learning and machine learning methodologies.
Confidential
Big Data / BI Architect
- Met with Business Owners to understand their Business requirements in order to develop new global supply chain business intelligence solution.
- Designed and implemented real-time Spark streaming applications integrated with Kafka to handle large volume and velocity data streams in a scalable, reliable and fault tolerant manner; implemented Kafka offset management and application monitoring framework to monitor data inflow using Scala, Python, Spark streaming, and Kafka.
- Designed and implemented Big Data analytics platform for handling data ingestion, compression, transformation and analysis of 30+ TB manufacturing and sales data.
- Implement automation, traceability, and transparency for every step of the process to build trust in data and streamline data science efforts using Python, Java, Hadoop streaming, Spark, Spark SQL, Scala, Hive, and Pig.
- Designed and developed HIVE data warehouse on top of HDFS using Avro and ORC data format.
- Designed and implemented Apache Solr search engine solution at Amazon AWS to provide search and analytics capabilities across structured, semi-structured and unstructured datasets. The solution initially implemented for partner search complaint to complex hierarchical data access control list.
- Planned and implemented the Migration process for Business Objects XI 3.1 to 4.1.
- Developed new ETL mappings and processes to feed aggregated tables having global and regional KPIs.
- Successfully Implemented Global operational KPI project with limited resources and time, which resulted in savings.
- Member of Big Data Analytics and Data Lake development team that created Data Lakes, and analyzed data patterns and predictive analysis.
- In order to improve performance, scalability, and memory usage of processing large volume of sales data, adopt Spark and Spark SQL to build JCI Benchmark cubes and populate cube annual and quarter metrics using Scala.
- Applied security features of Business Objects and Tableau like Row level, Report level & Object level security in the universe to secure sensitive data. Successfully completed modelling, design and development of Global Operations BI Dashboard initiative for Confidential .
- Developed self-service interactive dashboards for usage JCI executive board by replacing multiple standalone SAP BO reports and Charts.
Confidential, san antonio, tx
Sr. BI Architect
- Collaborated with Business Users and Business Owners to understand their Business requirements and written technical specification and design documents.
- Created ticket in the JIRA for the tasks, created branches in the Git. Created data frames to ingest the hdfs files into hive internal / external tables with partitions.
- Worked on writing Sqoop wrappers to move data between Netezza and Hive tables.
- Leveraged best practices in BI implementation to efficiently use infrastructure and to produce successful project results.
- Applied security features of Business Objects like row level, report level & object level security in the universe so as to make the sensitive data secure.
- Created visualizations through SAP Dashboards 4.1 using several displays including tabular, bar, pie, trend, scatter and map.
- Worked on SAP Lumira and Design Studio for linking data sets and HANA connectivity for creating rich visualizations.
- Provided technical solutions, resolved issues with deliverables, performed technical reviews and guided the team.
- Handled production implementations and support of BO solutions, Administration of BI (Business Objects).
- Created new WebI reports by creating a universe using Oracle Data Source.
- Worked on the SQL tuning and optimization of the SAP Business Objects reports.
- Designed an executive level Dashboard that shows high level metrics. Dashboard included customers detail drilldowns, threshold gauges, alerting mechanisms a pop report.
- Developed Complex Universes by linking different data sources.
- Installed Business Objects Enterprise and Published reports on Business Objects Enterprise Server and got involved in Business Objects Enterprise Administration.
- Interacted with Business Sponsors and Business Analysts from Business to understand and gather requirements for new projects and enhancements.
Confidential, wilmington, de
BI Architect/Engineer
- Partnered with Data Architects to ensure the data cohesiveness and ensure appropriate performance tuning is in place.
- Provided Installation road map for Business Objects Enterprise and Published reports on Business Objects Enterprise Server and got involved in Business Objects Enterprise Administration.
- Implemented BI best practices with project teams and in turn saved project cost and time.
- Interacted with Business Sponsors and Business Analysts from Business to understand and gather requirements for new projects and enhancements. Updated the project status to Stakeholders on regular basis as per the project plan.
- Created projects, project deliverables and assigned the work tasks to the team in order to stay on budget.
- Prepared the MSR (Monthly) and MTR (Quarterly) reports and submitting them to SQA review.
- Represented the BI Operations team at Business Change and Technical forums.
- Developed Complex Universes by linking different data sources.
- Utilized Tableau Desktop 8.X to analyze and obtain insights into large data sets.
- Worked Tableau Server 8.X to publish the dashboards and setup the security for interactive dashboards.
- Configured and enabled the SSO using the Windows AD Authentication.
- Responsible for the execution of the work requests and ensure the on time delivery to Validation and Production environments.
