Big Data Architect And Principal Consultant Resume
SUMMARY
- Overall 20 years of experience in IT.
- Hands on Experience of working in Big Data technologies - Hadoop, Hive, Spark (Python and Scala), Kafka on Hortonworks, EMR and HDInsight. Can build from scratch without any vendor.
- Hands on experience and expert knowledge in NoSQL technologies - Cassandra, Hbase, MongoDB and Redis. Wrote books on Cassandra and available on Amazon.com.
- Solid experience in traditional databases - SQL Server, Oracle from design, development and administration.
- Solid experience in Data warehousing, Data mining and Analytics.
- Solid experience and advance level knowledge in Java, Scala, Python, R and Node.JS besides .NET technologies
- Heavily worked on various cloud technologies - AWS, Azure and IBM softlayer. AWS EC2, S3, EMR, SQS, SNS, Kinesis, Redshift, Dynamo, RDS, Data Pipeline and so on. Experience in Azure HDInsight and other Microsoft cloud technologies.
- Did Architecture and developed IOT and Mobile wallet projects using Cassandra and Spark.
- Designed and managed the development of Store order processing system using Node.js, Kafka, Spark Streaming and Cassandra.
- Designed and Built data lakes using Hadoop/Hive, Spark and NoSQL(Cassandra/HBase).
- Designed and Built high transactional mission critical databases using Cassandra/Hbase.
- Experience of managing projects with Waterfall, Iterative and Agile methodologies. Plan task and milestones, manage deliverables to timeline and manage team. Capable of identifying risks, issues and mitigate and priorities and focusing on them deliver required results with minimum direction and supervision.
- Have solid experience in Data Mining, Machine learning and Forecasting. Used Spark MLib, Python, R, SAS and so on.
- Can design and build Graph databases in DSE graph using Cassandra, TinkerPop and Gremlin.
- Can design and develop Micro services and Restful APIs with Node.js, GoLang and can extend to Python, Java and Scala.
- Created Write intensive Database as Service product using Kafka/Spark, Cassandra and Scylladb.
- Created Write and Analytic intensive Database As Service using HBase and Kylin using Hive and Kafka as source data systems.
TECHNICAL SKILLS
RDBMS: SQL Server 4.2.1/6.5/2000/2005/2008/ R2/2012/2014, Oracle, 9i/10g/11g/12c, MySQL
Languages: Visual Basic, .Net, ASP.Net, C#, C, C++, Java
Data Warehousing/BI: SSAS, MDX, Excel, SSRS, SharePoint, MS Power view, PowerPivot, Tableau
Data Modeling: Erwin, ER Studio
Big Data: Hadoop, Spark/Scala, Hive, Kylin, Kudos
NOSQL: Cassandra, MongoDB, HBase/Accumulo, Redis, ScyllaDB
Machine Learning: R, Python, Azure Machine Learning, SSAS 2008/2012/2014 , Spark Mlib, SAS
Reporting (Analysis and Forecasting): Powerview, Excel, Microstrategy, Tableau, Sci2, Pajek, Gephi, D3
Operating System: Windows, UNIX, Redhat/Oracle Enterprise Linux 5/6/7, Ubuntu Linux 12/14, CentOS Linux 6/7
ETL Tools: SSIS 2005/ 2008/2012/2014 , Informatica
Cloud: Amazon EC2, Microsoft Azure, IBM Softlayer
DevOps: Chef/Jenkins, Grafana, Prometheus
PROFESSIONAL EXPERIENCE
Confidential
Big Data Architect and Principal consultant
Responsibilities:
- Data model design and implementation using Cassandra.
- Installs and Setup several Cassandra clusters (some of them are multi-data center). Built on AWS EC2, IBM softlayer and also on Private DCs. Administrative tasks using node tool utility. Backups, security and Maintenance. Used TDE and role based security.
- Worked with Datastax and also with Apache open source. Brought Apache open source Cassandra to Datastax level by building monitoring tools with Grafan/Prometheus, Integrating Spark and so on. Created POC with ScyllaDB for better write performance.
- Designed and built Store order fulfilment system using Java, Kafka, Spark and Cassandra. IOT model.
- Setup various MongoDB clusters with Sharding and Replication.
- Setup MongoDB backups, performance tuning and Data modelling. Administering and Supports MongoDB clusters. Created collections, indexes and optimized queries.
- Upgrades existing MongoDB clusters to 3.2 from 2.6 and upgraded to WiredTiger storage engine.
- Leverages on CDC (SQL Server) and GoldenGate (Oracle) to provide real-time analytics.
- Designed and Created Level 2 aggregated tables using Spark in Cassandra for easy access from Tableau.
- ETL using Spark, Sqoop, Cassandra loader and SSIS as well using Python programs.
- Integrated Cassandra with SOLR and Spark.
- Worked with HortonWorks to build 25 node HBase cluster and worked with developers to create schema and so on. Written python scripts to implement backups. Implemented UAL security. Setup Ambari for monitoring cluster. Configured Phoenix and Squirrel client to access and manipulate data from HBase using SQL.
- Designed and created Inverted Index for Enterprise search using MongoDB. Installed and configured Redis for caching login credentials.
- Designed technical architecture for AAA BigData platform and Enterprise Datalake.
- Built Horton Works (HDP 2.5) 60 node cluster on my own without involving Horton works consultants. Worked with Hive in ACID and llap mode for better throughput and flexibility.
- Configured cluster with Kerberos security. Ranger configuration. Implements Role based security.
- Configured Hive, Spark (2.0/1.6), HBase, and Kafka, Ambari views, Zeppelin and so on.
- Helped development team in setting up standards for their data lake design.
- Converts some Hive tables onto HBase for faster access and ACID transactions.
- Programs in Spark/Scala for processing data and generating presentation layer.
- Designed Data model for ERS (emergency road service) call volume. The raw data has over billions of records. This is to analyze agent performance, future forecasting and reduce overtime costs.
- Wrote Microservices and data ingestion services for data migration from source databases like Oracle, Teradata, SQL Server and also from non-traditional databases.
- Code and deployments uses Chef, Jenkin, Subversion. System monitoring tools like Nagios, Xymon.
Confidential, Santa Ana, CA
Sr. Database Architect Lead (BI AND DW)
Responsibilities:
- BI Architecture Design for FAMS division. Identify Software tools to be used.
- Used Cassandra and MongoDB to build various Data Warehouse databases. Hadoop and Map Reduce (along with Pig and Hive) to process unstructured data.
- Data Warehouse Design used Star and Snowflake schemas. Physical and Logical Data modelling using Erwin. Used Cassandra to store data.
- Created POC using Mongodb and Hadoop for Title search project.
- Productionized MongoDB cluster and administration.
- Designed SSIS packages to extract the data from Relational and Non-Relational sources and populate newly designed data warehouse databases. Used Spark and Sqoop for data migration.
- Designed OLAP Cubes. Developed Cubes using SSAS 2012.
- MDX support for Report querying. KPIs. Writing and tuning T-SQL and Stored Procedures.
- Scheduled SSIS packages, processing cubes, backups and so on.
- Integrated R and Python modules for Forecasting usage.
- Database Maintenance, Administration and Support.
Confidential, Irvine, CA
Sr. Database Architect
Responsibilities:
- Designed and Architecture new Databases, Database server environments and New Data centers.
- Created Database maintenance jobs for backups, re-indexing.
- Designed and implemented data warehouse databases using SQL Server, Oracle and NOSQL databases (MongoDB and Cassandra).
- Installed and Setup SSIS, SSRS and SSAS for Business Intelligence team. Design ETLs and Data Models for new data warehouse databases and supporting them.
- Built SSAS cubes and wrote MDX queries to access data from cubes.
Confidential, San Bernardino, CA
Sr. Database Administrator
Responsibilities:
- Data modeling and schema design for new projects and in re-engineering existing projects.
- Worked with vendors in implementing Red Prairie (time clock), ACI (Debit and Credit card processing) and AMS (Coupons).
- Led Data Warehouse and BI team. Did UDM for new Data Warehouse. Helped developers in developing SSIS packages using SQL Server 2008 and Implemented. The Data Warehouse is over 1TB and contains over 500 million records in one of the fact table.
- Built Cubes using SSAS 2008 and transferring knowledge to Developers. Implemented Microsoft Parallel Data Warehouse with Microsoft SQL Server 2008R2.
- Helped developers in building Pivot table reports using Excel 2003/2010.
- Did Data Mining for Market Basket Analysis using Association Rules.
- Used various Data Mining algorithms available in SSAS for Clustering, Forecasting and so on.
- Set up linked servers and connectivity using different drivers to other SQL Server, Cubes, Oracle and other data sources for Data warehouse and for some OLTP systems.
- Pharmacy Coupon and Generic Drug analysis and Forecasting for their Pharmacies.
- Product seasonality forecasting.
Confidential, Huntington Beach, CA
Sr. Database Administrator
Responsibilities:
- 24 * 7 production support.
- Data Modeling for IT Grid and Data Warehouse project using ERWIN.
- Performance Tuned. Identified bottlenecks, proposed solutions and implemented them in production environment. Improved performance of procedures and SQL queries through tuning Indexes, Partitioning, and file groups and so on. Used Index Tuning Advisor and Execution Plans.
- Wrote stored procedures and T-SQL.
- Server Installation for 32-bit and 64-bit in Cluster Environment and Administration. Setting up Database Mirroring for some servers. Hardware design, RAID.
- DTS and SSIS Packages. Writing new packages, migrating from DTS to SSIS.
- SQL Server migration from 2000 to 2005.
- SQL Server 2008 Installation and setting up Development Environment.
- Setup and Administered Service broker and Notification services.
- Setup Log Shipping and Replication for Reporting Servers and DR.
- Used Traces and Profilers. Resolving Deadlock and other issues.
- Backed up and restored using Tivoli and Litespeed software.
- Designed and Implemented Role based Security. Data Encryption for securing sensitive information.
- Created and maintained Analysis services OLAP Cubes. Installing and supporting reporting services and Dynamic reporting using Office Web components.
- XML parsed, XQuery, XPath and XMLx and XMLDOM.
Confidential, Costa Mesa, CA
Database Administrator/Architect
Responsibilities:
- Coordinated physical changes to data bases.
- Coded, tested and implemented physical data base.
- Applied knowledge of data base management system.
- Designed logical and physical data bases.
- Mentored and knowledge transfer to other team members.
- Production database support.
- Hardware designed and changed for new systems and to improve performance for existing systems.
- Created and maintain multiple databases.
- Security, backups and server maintenance.
- Daily data migration jobs using DTS and Informatica.
- Wrote stored procedures and T-SQL, XML queries.
- Executed database change management requests.
- Tuned queries and stored procedures using Profiler and Query Analyzer.
- Set up disastrous recovery systems using log shipping.
- Set up replication environment.
- Provided logic to store PDF and Excel files in BLOB columns of database.
- Production database support.
- Installed Linux and set up environments.
- Installed Oracle Instances, define users and grant authorizations.
- Performed data transfer from SQLServer and mainframes using Informatica (ETL).
- Regular monitor and support.
- Built QA and development environments.
- Performed RMAN Backups and Recovery.
- Security and performance tuning.
- Developed PL/SQL programming and tuning.
- This project provided analysis data reporting thus gives an insight into daily business for senior management. The users can drill down data and create their own reports using drag and drop options on the run. The modules developed so far include Brokers, Plan Sponsors and Heat.
- Proof of concept.
- Estimated time and cost required to accomplish project.
- Schema Design (Star and Snow-flake).
- ETL (Informatica 6.2 and 7.1.1) Sources include SQLServer, Oracle 9i, and text based mainframe file, excel and Access.
- Built Cubes used Analysis Manager, MDX.
- Analysis Reported used Office XP OWC10 and Actuate.
Confidential, San Francisco, CA
Consultant employee
Responsibilities:
- Database Design and Data Modelling for OLTP and OLAP systems.
- Administered Databases in production environment.
- Installed, upgraded and configured Database servers.
- Database development mainly using TSQL and PL/SQL.
- Performance tuning and Setting up High Availability and DR servers using clustering, Log shipping and so on.
- ETL development and maintenance using mainly DTS and SSIS.
- Built Cubes used SSAS for analysis. Worked on SSRS, MDX and other reporting tools like Excel for reporting purposes.
