We provide IT Staff Augmentation Services!

Aws/ Data Engineer Resume

4.00/5 (Submit Your Rating)

SUMMARY

  • Around 6+ years of working experience in AWS Platform development, data Analytics and worked with various Big Data tools
  • This includes gathering business requirements, analyzing cost, granting permissions, prioritizing tasks, development, validation, monitoring the ongoing activities and providing support in case of activity failures.
  • Apart from development, I thoroughly interact with business users in understanding their actual requirements to provide better solutions.
  • Handle multiple projects at once, but never lost the track on any of them. This made me improve my multi - tasking skills.
  • As part of heavy scope projects, the involvement of external team members always happens, for either access approvals or connectivity establishments or for Audit purposes. On acquiring soft skills and human-relation skills I was always able to effectively communicate and get the things done at the right time, without postponements.
  • Always an active team person to initiate any events or happenings among our team.
  • Attended several workshops conducted by the Clients to improve the awareness and cause for developing such mammoth projects.
  • Was recognized from Directorial level and received Appraisals in United Airline for the work that was contributed.
  • Initiated several enhancements required for developing a strong infrastructure for one of our Data-centers and was able to approve and achieve them throughout the team.
  • Was one of the very first team members to filter and trim required Audit logs AHF which reduced a great complexity
  • An Active member to provide support for United Airline’s Third-party logistics to debug and fix the Jobs which runs thrice a day.
  • Tools and Technologies which I’ve used and involved so far are as follows:

PROFESSIONAL EXPERIENCE

Confidential

AWS/ Data Engineer

Responsibilities:

  • Secured Apache Hue’s default settings by configuring Https protocol, with self-signed certificates, to make sure the data is encrypted in Transit and in Rest.
  • Developed Apache presto and Apache drill setups in AWS EMR (Elastic Map Reduce) cluster, to combine multiple databases like Mysql and Hive. This enables to compare results like joins and inserts on various data sources controlling through single platform.
  • Created AWS RDS (Relational database services) to work as Hive metastore and could combine 20 EMR cluster’s meta data into a single RDS, which avoids the data loss even by terminating the EMR.
  • Spin up the EMRs clusters from 30 to 50 nodes which are memory optimized such as R2, R4, X1 and X1e instances with autoscaling feature.
  • Hive Being the primary query engine of EMR, we’ve created external table schemas for the data that is being processed.
  • Mounted Local directory file path to Amazon S3 using S3fs fuse, to have KMS encryption enabled on the data reflecting in S3 buckets.
  • Involved in developing shell script, where the logs generated by the users are collected and stored in AWS S3 (Simple storage service) buckets. This includes the trace of all user activities and a good sign of security to identify cluster termination and to protect the data integrity.
  • Configured Apache Hue’s file browser and S3 browser operations and created User’s home directories. Noticed different Name node port number in Hue, when running on different protocols, and was able fix the default settings for their access and approval.
  • Applied Auto scaling techniques to scale in and scale out the instances with given Memory out of time. This helped in reducing the number of instances count when the cluster is not actively in use. This is applied by even considering Hive’s replication factor as 2 leaving minimum 5 instances running.
  • Worked on EMR Security Configurations, to store the self-signed certificates as well as KMS keys created into it. This makes to spin up a cluster in an ease manner without modifying permissions after the call.
  • Automated Shell script in the backend using nohup service, in collecting access logs and security logs generated in Hue console.
  • Triggered Put Method for Audit and Security log files residing in S3 bucket to Middle ware applications, using AWS Lambda with Restful API’s.
  • Configured Apache Hue console and it’s hive-site.xml property files.
  • Performed partitioning and Bucketing concepts in Apache Hive database, which improves the retrieval speed when someone performs a query.
  • Used AWS Code Commit Repository to store their programming logics and script and have them again to their new clusters.

Confidential

AWS Platform Developer/ Data Engineer

Responsibilities:

  • Developed Airflow which is a workflow management platform to schedule the Dags to copy the same records from S3 to Snowflake in three categories.
  • Enabled AWS Data Migration services to fetch the data from Source systems viz Oracle to copy the records which are either inserted or updated on the fly to S3 bucket.
  • Classified Data records into 4 categories viz low, medium, high and larger which have their own sub folders in S3 to keep a track based on data volume and importance.
  • During the file transfer process from S3 to Snowflake, the file count in Dynamo Db is compared with the distinct extract file records copied to Snowflake to make sure all the files are involved as part of data transfer.
  • Any Inserts, Deletes or Updates are always captured in the Archive Stage database first in Snowflake, which would update the final table with purely active results and have updates stay in Purge table for 6 months and has a copy in historical tables
  • Developed Cloud Formation scripts to create DMS tasks which involves to setup several variables and connections in the Airflow console.
  • Enabled to trigger Lambda Services on top of S3 bucket whenever the file drops from DMS. This Lambda would update the record count in Dynamo DB to have a count of number of files dropped in every group.
  • Although the new development was taken place on AWS platform, some of the teams still used Teradata for generating their Cognos reports or so. For this we enabled connections between s3 and Informatica servers to transfer the files for Informatica Server.
  • Compared data validation from Oracle to S3 and S3 to Snowflake and the complete Oracle to Snowflake validation to make sure we are at-least approximate to the results which are generated on the fly.
  • Developed a POC using Attunity for fetching files directly from Source Oracle system to Snowflake directly, but which involved a large licensing cost and processing charges.
  • Enabled both CDC and Full load operations at any cause of DMS task failures.
  • Used Elastic Container Service and Elastic Kubernetes service as part of Jenkins Job to enable any code changes in repository get reflected to services involved.
  • We also enabled to capture the logs generated in Informatica server for weekly validation of files transferred to them for auditing.

We'd love your feedback!