We provide IT Staff Augmentation Services!

Bigdata Developer Resume

5.00/5 (Submit Your Rating)

CA

SUMMARY:

  • Over 9+ years of experience in building distributed, scalable and complex applications using Hadoop Eco System, RDBMS, Spark and Informatica
  • Around 4 years of experience in Hadoop ecosystems: Hive, Sqoop, Hbase and Flume
  • 2+ years of experience in Spark using pyspark and scala
  • Expertise in Map - Reduce frameworks with structured & unstructured data
  • Well Versed in Apache Spark Framework - Data Frames, SQL, JSON, PARQUET files
  • Good knowledge of Hadoop Architecture and various components such as HDFS, Map Reduce, Hive, Job Tracker, Task Tracker, Name Node, Data Node and YARN
  • Expertise in SparkQL/HiveQL and implemented with best practices.
  • Worked on exporting/importing data using Sqoop from HDFS to Relational Database Systems and vice-versa
  • Proficiency in distributed query-processing tools like Hive, Spark SQL
  • Used tools like SQOOP, Kafka to ingest data into Hadoop.
  • Good understanding of cloud configuration in Amazon web services AWS
  • Excellent understanding of Tableau visualization and experience in building dashboards
  • Working knowledge in connecting Power Center to Teradata using FASTLOAD, BULKLOAD, MULTILOAD and TPUMP
  • Involved in full life cycle, which includes Data Analysis, Data Validation, Data Verification, Data Cleansing, Data Completeness and identifying data mismatch., and execution of the Process and Documentation.
  • Well versed with the concepts, design and development of DataMart s, Data Warehouse using Star Schema and Snowflake Schema and also implementing Decision Support System.
  • Excellent knowledge in database design, relational integrity constraints, OLAP, OLTP, Cubes and normalizations.
  • Extensive experience in DW Tools like Informatica, Oracle, TERADATA
  • Strong problem solving skills, quick learner and able to work independently as well as a team member of varying sized teams.
  • Ability to plan, manage, motivate and work efficiently as an independent or collaboratively in a team.
  • Self-motivated, enthusiastic and always keen to learn new methodologies and techniques

TECHNICAL SKILLS:

Big Data / Hadoop: HDFS, Map Reduce, Hive, Spark, Hbase, and Sqoop.

ETL Tools: Informatica 9.5.1 and Informatica Data Quality 8.6.1 , Teradata utilities

Reporting Tools: Tableau, Business Objects XI

Languages: PL/SQL, Python, Unix shell scripting

Database and Tools: Oracle 11g, Teradata V14, SQL Server 2007Oracle SQL developer 2.1

Operating Systems: Linux,UNIX and Windows

Other Tools: Eclipse, SBT,GIT

PROFESSIONAL EXPERIENCE:

BigData Developer

Confidential, CA

Responsibilities:

  • Built data pipelines to Load and transform large sets of structured, semi structured and unstructured data.
  • Imported data from HDFS into Hive using HiveQL
  • Involved in creating Hive tables, loading and analyzing data using hive queries 
  • Created Hive Partitioned and Bucketed tables to improve performance.
  • Developed a SQOOP Import Job, Shell Script & CRONJOB for importing data into HDFS
  • Used Tableau for visualization and building dashboards
  • To improve performance and optimization of the existing algorithms, explored different components like Spark Context, Spark-SQL, Data Frame, Pair RDD's, accumulators
  • Processed millions of records using Hadoop jobs
  • Implemented Spark code using Python for RDD transformations & actions in Spark application
  • Built reusable Hive UDF libraries for business requirements
  • Working with the leadership to understand scope, derive estimates, schedule, allocate work, manage tasks/projects, present status updates to IT and business leaders as required.
  • Define and contribute to development of standards, guidelines, design patterns and common development frameworks & components.

Data engineer

Confidential, CA

Responsibilities:

  • Translate ETL requirements into formalized designs and mapping documents.
  • Handled importing data from various data sources using Sqoop. Built pipelines to transform data using MapReduce/spark to load data into HDFS.
  • To improve performance and optimization of the existing algorithms, explored different components like Spark-SQL, Data Frame, Pair RDD's, accumulators
  • Implemented Spark code using Scala for RDD transformations & actions in Spark application
  • Hands-on experience with systems-building languages such as MapReduce in Python
  • Implemented hive tables on different file format like ORC and Parquet for efficient retrieval of needed columns
  • Optimized queries to use different types of joins like map side and bucket joins for efficient query retrievals
  • Involved in running Hadoop jobs for processing millions of records and compression techniques
  • Developed multiple MapReduce jobs for data cleaning and pre-processing 
  • Integrate external APIs (Salesforce, Google APIs, etc.) and ETL into data marts using Python
  • Built aggregates for canned reports using Hive
  • Implemented Broadcast & Accumulators variables in the application.
  • Developed Fast export scripts to spool files for third parties
  • Involvement in implementation of BTEQ and Bulk load jobs. 
  • Built ETL logic using Teradata bteq for marketing data models to measure campaign effectiveness
  • Created Different Fastload, Multiload and BTEQ Scripts to facilitate the ETL Processes 
  • TPT Execution is handled in Wrapper scripts and few occasions’s using Informatica loader’s to load data to different data bases for reporting purposes.
  • Expertise in performance tuning the user queries. execution of frequently used SQL operations and improve the performance. 
  • Perform unit test cases and run validation on the views to ensure data quality.
  • Created complex mappings using Sql and Flagged the record using update strategy for populating the desired slowly changing dimension tables.
  • Involved in discussions with Business Analysts and Business Users to design technical specification documents for Data warehouse
  • Followed Agile Methodology especially SCRUM software development process throughout development, QA and UAT
  • Leadership to advocate Dev team with User and other Technical Team Members.
  • Performed sizing for various EDW activities
  • Mentoring the developers and aiding the business and testing teams with the testing process. 
  • Audit architectural design and framework aka validation rule engine
  • To find data anomalies implemented scoring based audit.
  • To find data holes, created alerts if the data volumes are not in the desired threshold limits
  • Extensively researched and fixed error events /quality issues in production/QA environment.

DWH ETL Consultant

Confidential, CA

Responsibilities:

  • Worked in the full life cycle development of the warehouse from analysis through implementation which includes data modeling, Extract, Transform and Load process and Report development
  • Envisioned data model for future needs by keeping in mind the current needs and growth
  • Built Data quality screens in Informatica developer and designer to automate data cleansing in Data Warehouse
  • Followed Agile Methodology especially SCRUM software development process throughout development, QA and UAT
  • Strong business knowledge in Transportation domain especially Container Leasing business which involves Daily and Monthly Fleet Status, Daily Activity, Accounting - General Ledger, Revenue and Expense and Product Management
  • Responsible for Extraction, Transformation & Loading data in Tables using Teradata tools like bteq, multiload, fastload and Informatica Power Center 9.1/8.6 and Oracle PL/SQL.
  • Developed multiple Workflows for initial / historical data loading, using sessions, worklets, commands, decisions, events and timers
  • Designed and developed the ETL framework, including coordinating/executing.
  • Helped QA team to perform parallel testing with the existing system.
  • Worked closely with testers for Regression testing between cycles and releases.
  • Coordinating the “certification” and sign-off of the release.
  • Mentoring the developers and aiding the business and testing teams with the testing process. 
  • Used EXPLAIN PLAN to tune queries for better performance and also Extensive Usage of Indexes
  • Worked extensively on Teradata SQL query tuning using sliding window mechanism, temporary tables, collect stats, skew indexes and join indexes
  • Lead developer in 3NF database modeling and database optimization
  • Trained users on Business Objects especially in creating reports, variables, import external data sources, export to cms, access. etc.

We'd love your feedback!