We provide IT Staff Augmentation Services!

Sr. Big Data/hadoop Architect Resume

3.00/5 (Submit Your Rating)

Dallas, TX

SUMMARY:

  • More Seven years of comprehensive IT experience in Big Data domain with tools like Hadoop, Splunk, Hive and other open source tools/technologies in Banking, Finance and Cloud Computing.
  • Experience in Big Data Analytics with hands on experience in Data Extraction, Transformation, Loading and Data Analysis, Data Visualization using Cloudera Platform (Map Reduce, AWS, Hbase, HDFS, Hive, Pig, Sqoop, Flume, Hbase, Oozie, AWS, EME, Big Data, Hadoop, Spark, Spark, Lucene, SOLR, NoSQL, CouchDB, CouchBase, MongoDB, and Casandra & Splunk)
  • Experience working with JAVA, J2EE, JDBC, ODBC, JSP, Java Eclipse, Java Beans, EJB, Servlets, MS SQL Server.
  • Hands - on Experience on Cloud Platforms in OpenStack, OpenShift, CloudFoundry, Azure, AWS & Heroku.
  • Excellent understanding of Hadoop architecture and various components such as HDFS, YARN, Name Node, Data Node and Map Reduce programming paradigm.
  • Proficient in writing Hive SQL queries and Pig Latin scripts.
  • Strong experience in optimization/performance tuning of MR, PIG & Hive Queries.
  • Hands on experience in installing, configuring and using Apache Hadoop ecosystems such as Map Reduce, HIVE, PIG, SQOOP, FLUME and OOZIE.
  • Experience in building Big Data solutions using Lambda Architecture using Cloudera distribution of Hadoop, Kafka, Storm, Falcon, Trident, MapReduce, Cascading, HIVE, PIG and Sqoop.
  • Experienced in Worked on NoSQL databases - Hbase, Cassandra & MongoDB, database performance tuning & data modeling.
  • Experience in writing Ad-hoc Queries for moving data from HDFS to HIVE and analyzing the datausing HIVE QL.
  • Strong knowledge of Spark for handling large data processing in streaming process along with Scala.
  • Expertise in using Zoo-Keeper and Oozie Operational Services for coordinating the cluster and designing and scheduling workflows.
  • Familiar with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS), Teradata and vice versa.
  • Substantial experience writing MapReduce jobs in Java, PIG, Flume, Zookeeper and Hive and Storm.
  • Experience in development of Big Data projects using Hadoop, Hive, Splunk, HDP, PIG, Flume, Storm and MapReduce open source tools/technologies.
  • Extensive Knowledge on automation tools such as Puppet and Chef.
  • Hands on experience in installing, configuring and using ecosystem components like Hadoop MapReduce, HDFS, RedShift, Hbase, AVRO, Zoo Keeper, Oozie, Hive, HDP, Cassandra, Sqoop, PIG, Flume.
  • Extensive experience in SQL and NoSQL development.
  • In-depth understanding of Data Structure and Algorithms.
  • Experience in web-based languages such as HTML, CSS, PHP, XML and other web methodologies including Web Services and SOAP.
  • Extensive experience in all the phases of the software development lifecycle (SDLC).
  • Experience in deploying applications in heterogeneous Application Servers TOMCAT, Web Logic and Oracle Application. Server.
  • Extensive knowledge of NoSQL databases such as HBase.
  • Worked on Multi Clustered environment and setting up Cloudera Hadoop echo System.
  • Background with traditional databases such as Oracle, Teradata, Netezza, SQL Server, ETL tools / processes and data warehousing architectures.
  • Has handled multiple roles - Solution Architect, SOA Architect, Integration Architect, Application Architect, Systems Analyst, Developer, Programmer etc.
  • Has good understanding of integration technology, SOA patterns. J2EE & IBM Standard Methodology etc. Also have understanding of Big Data technology like Hadoop, Pig, Hive, NoSQL, Sqoop, IBM Big Insight, Spark.

TECHNICAL SKILLS:

Proficient in: HTML, SQL, PL/SQL, XML, MDX, UML, JavaScript. Terra data

Familiar with: Java, VB Script, Ajax, CSS and SOAP, C/C++

PROFESSIONAL EXPERIENCE:

Confidential,Dallas, TX

Sr. Big Data/Hadoop Architect

Responsibilities:
  • Delivered Working Widget Software using EXTJS4, HTML5, RESTFUL Web services, JSON Store, Linux, Hadoop, ZOOKEEPER, NO SQL databases, JAVA, SPRING Security, JBOSS Application Server for Big Data analytics.
  • Worked on the fast Access was developed using Confidential Redshift and finally connecting it to Qlikview for Visualizations and reporting.
  • Developed a custom AVRO Framework capable of solving small files problem in Hadoop and also extended PIG and Hive tools to work with it.
  • Working on use cases, data requirements and business value for implementing a Big Data Analytics platform.
  • Working on configuring and Maintaining Hadoop environment on AWS.
  • Developed application component interacting with MongoDB.
  • Working on Modifying Chef Recipes used to configure the Hadoop stack.
  • Involved in storing the MapReduce program output in Confidential S3 and developed a script to move the data to RedShift for generating a dashboard using QlikView.
  • Evaluate Puppet framework and tools to automate the cloud deployment and operations.
  • Working on Installing and configuring Hive, HDP, PIG, Sqoop, Flume, Storm and Oozie on the Hadoop cluster.
  • Created CloudWatch, AWS Data pipelines, Kinesis Realtime, Lambda, Firehose, API Gateway, BeanStalk, ECS & Spark Streaming projects on AWS Cloud environment.
  • Working in analyzing data using Hive, PIG, Storm and custom MapReduce programs in Java.
  • Developed Use cases and Technical prototyping for implementing PIG, HDP, HIVE and HBASE.
  • Working as a Big Data/Hadoop Architect on Integration and Analytics based on Hadoop, SOLR and web Methods technologies.
  • Working in implementing Hadoop with the AWS EC2 system using a few instances in gathering and analyzing data log files.
  • Working on data using Sqoop from HDFS to Relational Database Systems and vice-versa.
  • Working on loading files to HIVE and HDFS from MongoDB.
  • Founded and developed environmental search engine engine using PHP5, JAVA, Lucene/SOLR, Apache and MYSQL.
  • Used data science techniques to clean, enhance, and get the most popular topics. Made various views on AWS redshift database.
  • Led the evaluation of Big Data software like Splunk, Hadoop for augmenting the warehouse, identified use cases and led Big Data Analytics solution development for Customer Insights and Customer Engagement teams.
  • Worked on Distributed/Cloud Computing (MapReduce/Hadoop, PIG, HBase, AVRO, Zookeeper, etc.), Confidential Web Services (S3, EC2, EMR, etc.), Oracle SQL Performance Tuning and ETL, Java 2 Enterprise and Web Development.
  • Designed techniques and wrote effective and successful programs in JAVA, Linux shell scripting to push the large data including the Text and Byte type of data to successfully migrate to NO SQL Stores using various Data Parser techniques in addition to Map Reduce jobs.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and PIG jobs.
  • Worked on TOAD for Data Analysis, ETL/Informatica for data mapping and the data transformation between the source and the target database.
  • Working on Hive/Hbase vs RDBMS, imported data to Hive, HDP created tables, partitions, indexes, views, queries and reports for BI data analysis.
  • Developing data pipeline using Flume, Sqoop, PIG and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
  • Currently working on XML parsing using PIG, Hive, HDP and Redshift.
  • Working on architected solutions that process massive amounts of data on corporate and AWS cloud based servers.
  • Uses Splunk to detect any malicious activity against webservers.
  • Tuned the Hadoop Clusters and Monitored for the memory management and for the Map Reduce jobs, to enable healthy operation of Map Reduce jobs to push the data from SQL to No SQL store.
  • Working in importing streaming logs and aggregating the data to HDFS through Flume.

Confidential,New York City,NY

Big Data/Hadoop Architect

Responsibilities:
  • Worked on NoSQL databases including Hbase and MongoDB.
  • Used data stores included Accumulo/Hadoop and graph database.
  • Exploited Hadoop MySQL - Connector to store Map Reduce results in RDBMS.
  • Collected data from different databases (i.e. Teradata, Oracle, MySQL) to Hadoop
  • Used Oozie and Zookeeper for workflow scheduling and monitoring.
  • Worked on Designing and Developing ETL Workflows using Java for processing data in HDFS/Hbase using Oozie.
  • Working on AWSEC2 Instances Provisioning, AWSVPC setup, AWS Auto Scaling for availability of EC2 Instances and availability of applications
  • Created Hbase tables to store various data formats of PII data coming from different portfolios
  • Conduct vulnerability analyses; reviewing, analyzing and correlating threat data from available sources such as Splunk.
  • Designed, planned and delivered proof of concept and business function/division based implementation of Big Data roadmap and strategy project (Apache Hadoop stack with Tableau) using Hadoop.
  • Developed MapReduce jobs in Java for data cleaning and preprocessing.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Used Bash Shell Scripting, Sqoop, AVRO, Hive, HDP, Redshift, PIG and Java Map/Reduce daily to develop ETL, batch processing, and data storage functionality.
  • Responsible for developing data pipeline using flume, Sqoop and PIG to extract the data from weblogs and store in HDFS.
  • Working on extracting files from MongoDB through Sqoop and placed in HDFS and processed.
  • Worked on Hadoop installation & configuration of multiple nodes on AWS EC2 system.
  • Experienced in running Hadoop streaming jobs to process terabytes of XML format data.
  • Cluster coordination services through Zoo Keeper.
  • Involved in loading data from UNIX file system to HDFS
  • Installed and configured Hive and also written Hive UDFs.
  • Worked on setting up PIG, Hive, Redshift and Hbase on multiple nodes and developed using PIG, Hive, Hbase, MapReduce and Storm.
  • Installed and configured Hadoop through Confidential Web Services in cloud.
  • Design and implement data processing using AWS Data Pipeline.
  • Drove holistic tech transformation to Big Data platform to create strategy, define blueprint, design roadmap, build end-to-end stack, evaluate leading technology options, benchmark selected products, migrate products, reconstruct information architecture, introduce metadata management, leverage machine learning, productionize consolidated data store: Hadoop, MR, Hive, HDP and MapReduce.
  • Developed MapReduce application using Hadoop, Redshift, MapReduce programming and Hbase.
  • Developed Simple to complex MapReduce Jobs using Hive and PIG.
  • Worked on automate monitoring and optimizing large volume data transfer processes between Hadoop clusters and AWS.
  • Strong knowledge on Data Warehousing experience using Informatica Power Center.
  • Configure and manage Splunk Forwarders, Splunk Indexers and Splunk Search Heads.
  • Automated all the jobs for pulling data from FTP server to load data into Hive tables using Oozie workflows.

Confidential,Seattle,WA

Big Data/Hadoop Developer

Responsibilities:
  • Worked on analyzing Hadoop cluster and different Big Data analytic tools including PIG, Hbase database and Sqoop.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Built Analytics KPI engine using Python and PIG.
  • Worked on the data to import into AWS instance and Mapreduce jobs will be executed to analyze the data.
  • Developed fielded search dataset using Hadoop and Accumulo.
  • Working with Dell to scale their existing data pipeline to handle 10x the data to match our growth trajectory.
  • Worked in Big Data Analytics Initiative from system design to live operation in production.
  • Created customizations report using pentaho and MongoDB/ Cassandra provisioning/ decommission.
  • Installed Oozie workflow engine to run multiple Hive, Redshift and PIG Jobs.
  • Production and pre-production cloud infrastructure based on AWS cloud.
  • Worked on supporting and managing Chef Server.
  • Worked on Hadoop AVRO Files to Network File System for recording Audit data.
  • Worked on supporting MapReduce Programs those are running on the cluster
  • Help assisting finding bugs and learning Tableau Desktop for Splunk.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented a script to transmit sysprin information from Oracle to Hbase using Sqoop.
  • Implemented best income logic using PIG scripts and UDFs.
  • Worked on tuning the performance PIG queries.
  • Involved in loading data from UNIX file system to HDFS.
  • Worked on Apache web log data into Hadoop, transform it into a standard AVRO format, and output it to a Snappy compressed AVRO file.
  • Installed and Configured AWS Data Pipeline to run multiple AWS.
  • Expertise in developing and deploying Splunk, installing Splunk forwarders and developing dashboards.
  • Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
  • Built firmware check-in service using MongoDB, Java/Spring as both hosted and cloud service.
  • Develop reusable tools to efficiently handle the AWS migration pipeline, mostly scripts.
  • Installed Oozie workflow engine to run multiple Hive, HDP, Redshift and PIG jobs.
  • Supported in setting up QA environment and updating configurations for implementing scripts with PIG and Sqoop.

Confidential,New York City,NY

JAVA/J2EE Developer

Responsibilities:
  • Database analyzing, design and implementation.
  • Worked on entire data pipeline for automating using Flume and for the jobs scheduled periodically using Oozie.
  • Used JavaScript for client side validation.
  • Database connections and code implementation.
  • Used Python because supports multiple programming paradigms, including object-oriented, imperative and functional programming styles.
  • Expertise in writing shell scripts and Oozie workflows
  • Worked on developing and extending serialization frameworks like AVRO.
  • Development and maintenance of splunk dashboards based on the requirements.
  • Worked on configure AWS EMR (Elastic MapReduce).
  • Written the Apache PIG scripts to process the HDFS data.
  • Reviewing the existing Hadoop environment and make recommendations of new features that may be available and performance tuning with the other tools like Hive, PIG, MapReduce, Storm and Flume.
  • Worked on Monitoring, Replication and Sharding Techniques in MongoDB.
  • Performed Manual Testing, reported defects in JIRA and was responsible to keep track of them.
  • Used BI solution as a custom application build using OBIEE as front end and ODI and PL/SQL used for ETL. Developed HTML and JSP pages using Struts.
  • Designed GUI Components using Tiles frame work and Validation frame work.
  • Worked in Installation, Configuration and Management of Hadoop Cluster spanning multiple racks using automated tools like puppet and chef.
  • Worked on providing support for AWS Data Pipeline.
  • Worked on automating the jobs using Oozie in the project.
  • Used sequence and AVRO file formats and snappy compressions while storing data in HDFS.
  • Monitoring Hadoop scripts which take the input from HDFS and load the data into Hive.

We'd love your feedback!